Beyond Prediction Accuracy: Operational Tests for Physical Understanding in World Models
Abstract
World models are often credited with physical understanding from predictive accuracy or interpretable internal structure. We argue that neither is sufficient. We propose a six-stage Physical Representation Validity Ladder in which downstream claims are admissible only after prospectively specified checks of substrate competence, productive computation, representation fidelity/specificity, causal load-bearingness, and intervention-based repair. A sequence of synthetic-world experiments illustrates why these gates matter. Slot-plus-relation substrates contained the intended modules yet failed identity stability, selective relation dependence, or productive reconciliation. A 36-cell scaling diagnostic yielded precise compute exponents, but all reference predictors lost to persistence, making the scaling scientifically uninterpretable. Finally, a fully preregistered co-representation campaign formed seed-consistent, bidirectional and coupling-dependent repair, but failed frozen world-state fidelity, productive-refinement, and sham-validity gates; the campaign therefore closed as substrate construction failure and the higher-level hypothesis remained unadjudicated. The results support a conservative methodological claim: prediction, decodability, causal interaction, productive repair, and control validity answer different questions and should not be collapsed into a single claim of physical understanding. We offer the ladder, counterfactual intervention pairs, matched controls, and no-rescue decision rules as an evaluation discipline for physical world models and as a validation prerequisite for stronger tests of scientific-discovery agents that can progressively reopen representational commitments, reconcile ascending empirical constraints with descending reconstructions from more general constraints, and deepen abstraction only when reconciliation at a shallower level remains insufficient.