Embodied XAI Verifies the Agent: Tracing Hierarchical VLM–PPO Decisions via 4D Spatio-temporal Saliency
Suzan Halim
Abstract
Verifying an embodied agent means more than checking whether it reached the right object: an agent that succeeds for the wrong reason is a silent regression waiting to surface. Outcome checks cannot distinguish the two, and per-frame 2D saliency --- the standard process-level probe for vision models --- does not survive contact with a continuously operating robot, because every heatmap is anchored to a different, moving camera and therefore cannot be pooled, compared, or tested for repeatability. We contribute an auditing methodology that closes this gap. \emph{4D spatio-temporal saliency} (i) runs integrated gradients against the exact sign-inclusive token span of the velocity command the agent actually sampled while driving --- teacher-forced, over the same frame window and memory snapshot that conditioned the decision --- (ii) unprojects each pixel's attribution through depth and known camera pose into a shared world frame, and (iii) accumulates the resulting weighted point clouds across queries and trials into one heat field over the 3D arena that grows through time. Anchoring attribution to the environment rather than the camera makes it poolable, and pooling makes it verifiable: convergence sweeps certify a hot region as repeatable rather than single-trial noise, outcome-split fields isolate the attention signature of each choice and of failure, and time-resolved accumulation localizes the moment of target commitment --- or its absence. We instrument a Unitree Go2 quadruped whose frozen Qwen2.5-VL-3B planner emits continuous velocity commands to a 50\,Hz PPO locomotion policy and audit it under deliberately ambiguous instructions: across 300 trials a mirrored two-target probe with a color-neutral control exposes a dominant left-side bias (87\%, $p\approx1.3\times10^{-14}$) and a significant prompt-dependent color effect ($\chi^2=7.52$, $p=0.006$) on pixel-identical images, and the heat fields then localize where and when those biased decisions were grounded.
Chat is not available.
Successful Page Load