Where Does the Alignment Lie? What Language Models and Reading Brains Actually Share
Maria K Tzevelekou
Abstract
Language models can predict neural activity recorded during reading, and this is frequently interpreted as evidence that models and human brains process language in similar ways. However, prediction alone does not explain what is being shared. We therefore ask which model features are responsible for this predictive alignment, testing the question at the word level where semantic effects can be distinguished from sentence-level topic. We evaluate Pythia and Gemma-3 against word-level EEG recorded during natural reading (ZuCo 2.0), and use sparse autoencoders (SAEs) to decompose the observed alignment while applying confound controls, permutation tests, held-out validation, and a language-free baseline. Three main findings emerge. First, a position-and-word-rate baseline (PWR), containing no linguistic information, predicts EEG at least as well as every model we evaluate, including a contemporary 4B model. Second, the band with the strongest raw score, beta2 ($r=0.19$), loses 76\% of its score after confound control, indicating that the apparent effect primarily reflects word position rather than linguistic processing. Third, the features that remain after all controls are structural: they correspond to sentence boundaries and grammatical function words, while no semantic survivors are found up to 12B parameters. Within Gemma-3, this structural component increases with model scale before reaching a plateau, but it never becomes semantic and never outperforms PWR. Thus, word-level EEG-LLM alignment is genuine and interpretable, but its content is structural rather than semantic, and it does not surpass a language-free baseline even for modern models.
Chat is not available.
Successful Page Load