What Do SAE Features Encode? Evidence from Human Neural Activity
Abstract
Sparse Autoencoders (SAEs) decompose dense LLM activations into sparse, interpretable features. However, evaluating whether SAEs extract genuinely meaningful structure remains challenging. Current approaches rely on model-internal metrics such as reconstruction fidelity and automated LLM scoring, which assess SAEs as mathematical decompositions but provide no external validation that the extracted features correspond to anything outside the model. To address this issue, we propose a new validation approach: comparing SAE representations against human neural activity. The validation rests on a shared computational principle, since both SAEs and biological neural systems implement sparse coding over overcomplete populations. Through extensive experiments using EEG recordings of naturalistic reading and SAE features from three large language models, we demonstrate that SAE features systematically align with human brain activity. We further show that this alignment is dominantly carried by the SAE's learned sparse code, rather than by generic architectural properties or the scale of the training data. This work is the first attempt to validate pretrained SAE features against human brain activity, establishing biological alignment as a complementary benchmark for mechanistic interpretability research beyond model-internal metrics.