The Linear Representation Hypothesis for Vision-Language-Action Models
Abstract
The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on the internal representations of large language models. A growing body of work has extended this perspective to vision-language-action (VLA) models. We argue, however, that this extension is not straightforward. Unlike high-level concepts in LLMs, which are typically studied as properties of a given representation, a quantity of interest (QoI) in a VLA is coupled to the system's closed-loop dynamics. The representation influences the actions selected by the policy, which alter the environment and, in turn, the subsequent representation. Accordingly, we develop a rigorous, signature-based formulation of the LRH for VLA models that unifies both representations and policies. On the representation side, we show that VLA representations can encode sufficient information to predict the future evolution of a quantity of interest (QoI) under a candidate action trajectory. More specifically, we establish the existence of a finite-dimensional shared representation in which multiple propagated physical QoIs can be recovered to arbitrary accuracy via action-conditioned linear probes. On the policy side, we formulate the stochastic VLA policy as a signature generalized linear model. Leveraging the information-geometric structure of this formulation, we show that the expected future QoI varies monotonically along a linear path in natural parameter space, thereby enabling linear steering.