Decision-Level Auditing of Latent Robot World Models
Anikait Rana ⋅ Aditya R Udupi ⋅ Siddhant Rana
Abstract
Latent world models are often judged by prediction error, but robot planning depends on whether predicted futures rank candidate actions correctly. We introduce a decision-level audit that separates these questions. On frozen 300-action Push-T menus from released JEPA-WM and DINO-WM planners, we compare the action chosen from predicted terminal latents ($P$) with the action chosen after scoring simulator-realized endpoints in the same latent space ($O$). We evaluate both choices using goal-distance utility $G$ (negative final-state distance) and task reward $J$. Across 20 trajectory clusters and 562 model-matched rows, predicted scoring substantially outperformed uniform menu selection. Replacing predictions with realized endpoints, however, did not reliably improve physical outcomes: the raw $O$-versus-$P$ effect was $-5.30$ in $G$ (95% CI $[-9.64,-0.71]$, Holm-adjusted exact $p=.0755$) and $+0.042$ in $J$ ($[-0.176,0.240]$, $p=.701$). Across all 577 merge-valid diagnostic artifacts, $P$ and $O$ selected different actions in 337; only 48 and 38 of those switches reversed the strict $G$-best and $J$-best pairwise ordering, respectively. The central result is a separation: the latent scores were useful for choosing actions, but more faithful terminal latents did not reliably make those choices better. A post-hoc, prediction-only near-tie score nevertheless tracked these selector changes (exploratory AUROC $.714$, 95% trajectory-bootstrap CI $[.665,.762]$): the held-out top ambiguity quarter switched 82.4% of the time versus 51.0% otherwise. It did not identify higher planner regret, so we treat it only as a candidate trigger for prospective verification rather than call it physical risk or epistemic uncertainty.
Chat is not available.
Successful Page Load