Embodied XAI: Tracing Hierarchical Agent Decisions via 4D Spatio-Temporal Saliency
Abstract
Robots steered by vision-language models are being proposed for spaces people occupy, where an unaccountable choice carries real cost. Yet when such an agent silently picks a target, per-frame 2D saliency cannot say what in the physical world drove it: the planner is re-queried continuously while the robot drives, so each map sits in its own moving camera frame, leaving thousands of disconnected images that cannot be pooled, sliced, or compared. We present 4D spatiotemporal saliency, a model-agnostic pipeline that attributes the planner’s own sampled steering token, lifts each attributed pixel into world coordinates through depth and pose, and accumulates the result across queries and trials into one heat field growing through time. Anchoring attribution to the scene rather than the camera turns an explanation into an object that can be added up, grouped by outcome, differenced against a reference, replayed, and physically intervened upon — operations a stack of camera-frame heatmaps does not support at all. We instantiate it on a quadruped driven by ambiguous instructions, and use the field to certify attention patterns as repeatable, separate the signatures of success and failure, watch a decision harden, and run closed-loop occlusion experiments no image-coordinate mask can express.