Which Latent Coordinates Drive World-Model Dynamics?
Abstract
Standard probing techniques reveal which physical variables are encoded or "readable" within a world model's latent state, but they fail to identify the causal mathematical coordinates that actually drive its transition dynamics. To bridge this gap, we introduce an action-conditioned intervention test that discovers dynamically used coordinates without requiring ground-truth state labels. From a single latent state, we compare the predictive rollout of a perturbed action against the rollout of a geometrically transformed latent state. If the mathematical transformation successfully reproduces the model's response to the physical action perturbation—judged against a variance-matched null baseline—we conclude the model actively uses that coordinate for prediction. Applying this framework to a simulated two-joint robotic arm demonstrates a strict dissociation between readability and dynamic use. While standard probes show both joint angles are strongly decodable from the model's memory, our test proves the transition model relies almost exclusively on the proximal joint, yielding a conjugacy margin six times larger than the distal joint. Furthermore, by holding the physical action space and joint dynamics constant while altering the visual rendering, we demonstrate the mechanism behind this coordinate formation: predictive training organizes latent dynamics entirely around observation geometry. Thus, the world model constructs its internal physics engine based strictly on the variable distinctions available in the visual input.