Auditing World Action Models Through Imagined-Future Interventions
Abstract
As embodied agents move toward open-ended deployment, plausible predicted futures can create confidence that an action is grounded even when the action may be computed through another route. World action models are a concrete safety and oversight case: they produce predicted futures alongside robot actions, but co-generation alone does not causally establish how the future affects the action. We propose a causal audit with two intervention stages. Stage 1 measures how strongly and in what direction future content changes behavior, and Stage 2 identifies the internal pathway carrying that effect. Across 22 states and six RoboLab tasks, donor future targets redirect Cosmos 3 actions and executed endpoints almost completely toward the donor rollout. Cosmos Policy shows a smaller first-decision-state effect across ten LIBERO-Long tasks. In Stage 2, restoring recipient K/V at predicted-future token positions removes 83–88% of Cosmos 3 steering. The audit distinguishes visual plausibility from action grounding, identifies model-specific reliance on future content, and exposes residual routes that matter for oversight of reliable embodied agents.