World Models Need Not Be Right to Rank Policies Right: Relative Bounds and Audit Limits
Ahanaf Ariq ⋅ Hasaan Ahmad
Abstract
World models can be inaccurate in absolute return yet useful for selecting policies when their errors are shared. We formalize this distinction through an exact pairwise simulation identity and a relative bound separating common-mode Bellman residuals from policy-specific error. This analysis motivates RAVEL, a selective policy tournament whose eliminations are grounded in held-out real outcomes. We evaluate 45 SAC, TD3, and PPO policies per task on DMC cheetah-run and walker-walk. At rollout horizon 500, model rankings remain strong, with Kendall correlations of $\tau=0.962$ and $\tau=0.800$. Same-run differential-error ratios are $0.674$ and $0.106$, showing that relative error can be substantially smaller than two absolute errors. However, mean paired audit-error correlations are only $0.007$ and $0.090$, while pairing changes standard deviation by $-2.5\%$ and $+1.9\%$. Simultaneous intervals achieve $99.98\%$ and $99.75\%$ empirical coverage but certify only $13.2\%$ and $0\%$ of decisions at eight audits. RAVEL improves low-budget walker-walk regret, although real-only successive halving performs better at the largest budget and the precommitted twofold efficiency target is rejected. These results demonstrate that absolute calibration, ranking fidelity, and audit utility are distinct notions of world-model correctness.
Chat is not available.
Successful Page Load