Distinct Patterns of Brain Alignment of Language Models across Linguistic Domains
Przemek Kubiak ⋅ Suchir Salhan ⋅ Paula Buttery ⋅ FERMIN MOSCOSO DEL PRADO
Abstract
Language models can reproduce striking aspects of human language behaviour, yet it remains unclear whether the representations they learn develop along trajectories resembling those of the developing brain. We ask when, and in which linguistic domains, brain–model alignment emerges during language acquisition. We compare representational geometry from four developmental fMRI datasets spanning ages 5–15 with hidden-state representations from language models across training, including models trained on developmentally realistic data budgets. Across semantic, phonological, grammatical, and plausibility domains, we find no single developmental trajectory. Grammatical representations show robust positive brain alignment throughout training, whereas semantic and plausibility alignment reverses sign as models acquire more data; phonological alignment remains largely indistinguishable from zero. Strikingly, plausibility provides a direct dissociation between cognitive behaviour and neural representation: models become dramatically better at plausibility judgments while becoming less brain-aligned, with accuracy and alignment correlated at $\rho=-0.64$. This relationship is substantially stronger than any corresponding effect in the other domains. Moreover, neither circuit sparsity nor phenomenon selectivity tracks behavioural accuracy or brain alignment. These results challenge the view that better language performance, more brain-like representations, and increasingly modular circuits necessarily emerge together. Instead, brain–model correspondence is domain-specific and dynamically reshaped by training, suggesting that behavioural competence alone is an insufficient proxy for developmental or neural alignment.
Chat is not available.
Successful Page Load