Verbalizable Representations emerge before Workspace Functions in Language Models
Abstract
The Jacobian Lens identifies vector representations that encode the model’s potential to produce particular tokens in future outputs. Together, the vectors define the Jacobian space (J-space), a proposed workspace for these verbalizable representations. It remains unclear, however, when these representations first become recoverable through the J-lens and how their causal influence on model behavior develops throughout pretraining. We study J-space across Pythia scales and pretraining checkpoints and measure representation availability through J-lens readouts, causal relevance through targeted ablations, and workspace-like functional properties through intervention-based tests. We find that representational availability appears early for information from the current and earlier context, while internally derived representations become available gradually. Causal relevance develops later, but workspace-like functional properties remain weak throughout training. These results show that availability, causal relevance, and function follow distinct developmental timelines.