Verbalizable Representations Emerge Before Workspace Functions in Language Models
Arjun Dabir ⋅ Ailun Shen ⋅ Michael Chan ⋅ Bangzishu Huang ⋅ Shanduojiao Jiang
Abstract
The J-Space provides a way to separate when internal representations become readable from when they begin to affect model behavior. We study its development across various Pythia model training and scale. Information from the current and earlier context becomes readable early, while latent intermediate representations emerge more gradually. J-Space ablations increasingly affect model performance later in training, indicating growing causal relevance. However, stronger functional properties remain weak. These results suggest that readability, causal relevance, and functional control mature along different timelines.
Chat is not available.
Successful Page Load