Prediction Fidelity Is Not Decision Fidelity: A Counterfactual Action-Ranking Diagnostic for Robot World Models
Abstract
World models increasingly score imagined robot futures, yet trajectory error does not directly test the quantity a candidate-based planner uses: the ordering of candidate actions. We introduce a controlled counterfactual action-ranking diagnostic for decision fidelity, measuring whether predicted futures preserve the oracle order over a fixed candidate action bank from a matched query state. In a state-based manipulation task, an unobserved actuation-response factor is inferable only from a four-step interaction context. Changing it alters the oracle top-ranked action in 96.5% of held-out states and reverses 45.7% of comparable action pairs under the strongest shift. As a controlled intervention, we hold architecture, data, optimizer, and state-prediction target fixed and add a lightweight pairwise ranking objective computed solely from predicted future states. Across three paired seeds, ranking-aware training incurs a 2.0% increase in trajectory RMSE (4.009 versus 3.929 mm), while improving pairwise ranking accuracy and top-1 action selection by 1.67 and 6.81 percentage points, respectively, and reducing mean decision regret by 23.7%. The effect persists after an 11.1× capacity increase. These results expose an evaluation failure mode missed by aggregate prediction error and motivate reporting action-order accuracy and regret whenever imagined rollouts rank robot actions.