Partial Composition, Not Blind Substitution: Decomposing a Grounding Failure in Vision-Language-Action Models
Abstract
Robots controlled by vision-language-action models are given their tasks in ordinary English, and the instruction is the only channel through which a user can change what such a robot does. These models now perform well on standard manipulation benchmarks, yet a growing body of evidence finds that they lean on visual regularities rather than the words: they take shortcuts on the instruction's most salient factors, and they fail on object and place combinations they were never trained on. What that evidence does not settle is why, because several different mechanisms predict the same aggregate failure rate. We test those explanations against each other using matched controls, and ask whether standard interpretability methods can locate the failure inside the network. Using a LIBERO-finetuned SmolVLA, we show the policy identical camera images paired with instructions that differ in exactly one word. The failure is narrower than a "vision overrides language" account predicts. When the arm is already carrying an object toward one goal and the command names a different one, the policy follows the command on 94.6% of trials [91.6, 97.4], and on 74.5% [68.8, 80.1] of neutral states. Failure appears only for untrained word combinations, and the resulting errors are structured: they favour places the object was previously paired with, by 11 and 10 percentage points above a uniform choice among the incorrect answers. The behaviour is nevertheless not a fixed object-conditioned destination lookup, because which place the policy heads toward still depends on which place the command names, selected on 34 to 67% of untrained trials. Paraphrasing the instruction leaves the effect in place on both checkpoints, though one of them shows sharp, destination-specific sensitivity to the exact wording, which we report as evidence that surface-form memorisation is a separable failure mode. We then report two negative results together with the controls that establish them. A 130-site causal patching sweep appears to localize the failure, until the identical sweep on a contrast containing no failure reproduces the same profile; no site survives correction. A linear probe finds the instruction encoded in the vision-language backbone, but its results inside the action expert reverse between checkpoints. We release the stimulus generator, a competence gate that finds only 2 of 8 public checkpoints usable, and seven measurement artifacts that produced convincing but incorrect results in our own pipeline. The gate and the artifacts point the same way: here a localization result was only as trustworthy as the control that could have overturned it, and we release the contrastive conditions such a control needs. The confound behind our sweep is architectural rather than incidental, since an action decoder reading its backbone through per-layer caches ties an intervention's reach to its depth; we state that as a hypothesis for other encoder-decoder policies rather than a demonstrated general result. Code, stimuli and all run artifacts will be made public upon acceptance.