Say the Same, Act Differently: Text-Orthogonal Action Subspaces in Reasoning Vision-Language-Action Models
Abstract
Vision-language-action (VLA) policies increasingly generate natural-language rationales before executing embodied actions. These rationales are attractive as monitors for safety-critical behavior because users or automated checkers can inspect whether the model appears to understand the scene and intend a safe action before execution. We show that this signal can be misleading. In many reasoning VLAs, the rationale generator and action head share continuous hidden states, while the generated text reveals only part of that representation. This creates a Say the Same, Act Differently failure mode, where rationale text remains stable while the action output changes substantially. We first study this phenomenon by introducing a new statistical method to separate text-sensitive and action-sensitive directions in hidden-state space of a reasoning VLA. Across the driving VLA Alpamayo and the manipulation VLA InstructVLA, we show that the dominant action direction lies largely outside the active text span, with mean projection ratios of only 1.9% and 5.3%. We define the residual as a text-neutral action-subspace (TNAS) direction, which tests whether action-relevant directions remain after removing measured text-sensitive directions. We then use TNAS to guide white-box pixel-space projected-gradient-descent (PGD) attacks on Alpamayo. TNAS-guided PGD successfully uses a localized two-forward-camera patch to induce a 30.5 m mean ADE shift in trajectory planning while preserving the exact rationale in 86.7% of cases. Moreover, across multiple text-preserving PGD objectives, the optimization direction consistently aligns with TNAS. These results reveal a fundamental shared-cache monitorability gap in reasoning VLAs: stable rationale text should not be treated as sufficient evidence that the action output remains stable.