Position: World Models of Physical Sites Must Discover Their Own State
Abstract
We argue that a world model of a specific physical site must discover its sufficient state from evidence rather than inherit it from sensors or designers. Video models treat pixels as state, entangling persistent variables with viewpoint, illumination, and occlusion while exposing force, vibration, material, and machine mode only indirectly. Digital twins begin from engineered variables that maintenance, wear, or intervention can invalidate. We call a model that avoids both defaults a Reality Model: it maintains a persistent predictive state tied to one site; accepts heterogeneous time-stamped sensors without letting one family define the world; estimates calibration and timing as state; and revises which variables the state contains when predictions fail against physical outcomes. Added sensors do not change a site's degrees of freedom, but may make them identifiable and reduce the ambiguity the model must carry. We state falsifiable hypotheses on state discovery, observability, identifiability, persistence, error attribution, adaptation, transfer, and agents, with longitudinal comparisons against video, geometric, object-centric, state-estimation, and hybrid baselines. The position fails if general pretrained models match persistence, calibration, and recovery under controlled observations and honest resource accounting, or if learned state discovery adds nothing over engineered state and system identification.