Predictable Before Action, Resistant to Correction: Failure Signals in Vision-Language-Action Models
Abstract
Long-horizon manipulation with vision-language-action (VLA) policies remains unreliable under distribution shift. We ask whether a policy’s memory state signals impending failure early enough to act on. On MemoryVLA, a linear probe on the fused cognitive state predicts eventual failure at 0.862 AUC before any episode could have ended, generalizing to unseen perturbation dimensions (0.847) and unseen scenes (0.811); most of this signal is present in the first observation. We test four interventions gated on this signal (additional sampling, a scripted recovery macro, slowed execution, and memory-bank clearing), each verified to change the robot’s behavior; on this policy and benchmark none reliably improves success, though an oracle shows headroom exists. Aborting predicted failures early saves 27% of compute at 2% cost in successes.