Reaction Without Attribution: Hijacking a VLA’s Actions Reveals No Linearly Readable Efference Copy
Ziwei Liu
Abstract
Biological motor systems carry an efference copy: a record of the outgoing command that lets the system attribute observed change to itself versus the world. We ask whether a post-trained vision--language--action (VLA) policy carries anything analogous. During LIBERO-Goal rollouts of OpenVLA-OFT (7B, SFT), we causally hijack the executed action chunk with probability $0.25$, primarily by swapping in another same-task episode's chunk, and probe hidden states (9 layers $\times$ 3 token pools) for the self-vs.-hijacked distinction. No hidden-state probe separates from a selection-aware shuffled-label floor ($0.536$; $n{=}1{,}716$), even when the probe is handed the commanded chunk or both pre- and post-action states. The null is not unrecoverability: explicit command--outcome comparison features decode the label at $0.69$--$0.875$ across intervention types, and the commanded chunk itself reads out of the pre-action state at $R^2{=}0.64$. Yet the policy reacts: next-call entropy rises and action log-probability falls (episode-level $p\!\approx\!10^{-3}$)---but only for visually novel interventions, and not for the most mechanically blatant one (freeze). SFT post-training thus yields a policy that compensates for hijacked actions through visual servoing while encoding no linearly readable attribution signal---a substrate-level account of "compensation without encoding,'' and a pre-registered baseline for asking what RL post-training adds. Code, data, and an append-only preregistration log will be released.
Chat is not available.
Successful Page Load