Where Do Runtime-Correction Gains on a Frozen VLA Come From? A Placebo-Controlled Dissection
Abstract
Runtime correction briefly takes control when a frozen vision-language-action (VLA) policy appears to fail. Its reported gains, however, can be confounded by the stochastic outcomes of the base policy. We introduce a placebo-controlled evaluation for runtime correction: paired intervention and base rollouts, a same-seed base re-run placebo, and a rescue-breakage decomposition. We use this evaluation to dissect a pipeline comprising a stall monitor, a privileged teacher, counterfactual data admission, and a distilled joint-space corrector on four RoboTwin simulation tasks. The analysis changes the interpretation of apparent rescues: on the only task with a detected effect, most naive recovery is explained by the placebo. A simple timer-triggered scripted retreat matches the learned pipeline in a preregistered replication, while neither intervention shows a detected benefit on the remaining tasks. Thus, under this deployed configuration, recoverability depends on the task rather than the sophistication of the repair. We provide a break-even screening rule and limit our claims to this simulator benchmark.