Where Structure Enters a World Model: From Predictive Fit to Physical Rollout
Abstract
Predicting the next frame or token does not guarantee that a model has learned a consistent rule for how a physical system evolves. We ask how added physical structure can turn predictive models into more reliable simulators while preserving their original learning tasks. We compare four places where structure can enter: training targets, temporal representations, pairwise interactions, and the transition used for rollout. In an image-first V-JEPA2 model, learned Hamiltonian and Lagrangian heads generate organized phase-space trajectories across held-out pendulum and spring initial conditions. Compared with latent autoregression, this structured path reduces simple-pendulum angle error by 89.6\% and spring position error by 80.7\%. We then use a three-seed planetary benchmark with 100,000 training trajectories to separate the effects of temporal and geometric structure. An equivariant pairwise adapter reduces autonomous rollout error from 1.805 to 0.875. Combining it with body-aware temporal learning lowers the error further to 0.498. Soft force and conservation targets improve adaptation from sparse force labels, with angular-momentum supervision reducing force-vector error by 76.4\%. Across both testbeds, where structure enters the model determines whether it improves physical readout, trajectory stability, coordinate robustness, or autonomous rollout.