How Do World Action Models Use Their Futures? Evidence from Mechanistic Interventions
Abstract
World action models generate future observations alongside robot actions, creating a potentially powerful, but poorly understood, control pathway in embodied AI systems. For safety-critical deployment, it is essential to determine whether these generated futures genuinely guide behavior, what information carries that influence, and where the influence enters the policy. Prior work shows that corrupting future features degrades performance, but destructive perturbations cannot distinguish causal control from generic distribution shift. We introduce reachable-donor future transplantation, which exchanges policy-generated futures between two feasible continuations from the same restored state while holding the recipient’s observation, instruction, and random draws fixed. This intervention yields a directional prediction: if generated futures control behavior, actions and simulator endpoints should shift toward the donor continuation. Donor futures nearly completely redirect Cosmos 3 actions and endpoints across 22 states and six RoboLab tasks; Cosmos Policy exhibits related early-state steering across all ten LIBERO-Long tasks. The safety-relevant control signal is model-specific: Cosmos Policy responds primarily to camera-visible robot motion, whereas neither robot-only nor object-only pixels reproduce Cosmos 3’s whole-future effect. Replacing donor-future attention keys and values with self-future counterparts removes 83-88% of Cosmos 3 steering in every tested state, identifying a dominant internal pathway. These results establish generated futures as causally active control signals rather than passive predictions. By exposing what information steers behavior and how it propagates through the policy, our framework supports more targeted auditing, monitoring, and intervention for reliable and controllable embodied AI.