The Score Is Not the Structure: Brain Alignment and Cross-Lingual Transfer
saman rahbar
Abstract
Researchers routinely claim that a model shares structure with something else, with the human brain or across languages, and support the claim with a similarity score. We ask a question that comes first: what does such a score read when the shared structure is absent, or when the tool used to measure it does not work? We run that check in two settings. In both, the score is not what it appears. In the first, a small classifier is trained to tell grammatical from ungrammatical sentences inside one language, then applied to another. Its accuracy falls as the two languages grow more different, which is the usual evidence that models encode structure common to both. But the classifier itself gets worse along the same axis. In four of our seventeen languages it performs no better than guessing, so $64$ of the $272$ language pairs are scored using a tool that does not work. Dropping those four halves the strength of the relationship. They are also the four most unusual languages, however, so dropping them shrinks the range of the predictor at the same time, and we show the two effects cannot be told apart here. The relationship is real; how strong it is cannot be recovered from this design. How the result is counted matters as much. The same data give a strongly significant $p$ of $0.0006$ when the $272$ pairs are treated as independent, and a null $p$ of $0.155$ when the seventeen languages are, which is the correct unit. In the second setting, a language model is trained to make its internal geometry resemble human brain responses, measured by a standard similarity score. The score rises from $0.09$ to $0.34$, against $0.54$ for the most that two independent groups of people agree with each other. But a model trained on a target whose correspondence to the brain has been destroyed still scores $0.31$. Only $0.028$ to $0.068$ of the rise is specific to the brain. Comparing the targets to each other, with no model at all, shows why so much is available for free: a destroyed target already sits at $0.204$ from the real one. Across the ranks we tested that figure tracks the target's dimension divided by the number of sentences, which suggests anyone scoring against a dimension-reduced target can estimate their own floor cheaply, and should check it. Finally we ask whether either correspondence is useful, and neither is. Steering a language along its own direction does move its grammatical preference, by $+6.2$ over a random direction in sixteen of seventeen languages, so the intervention plainly works and still shows no effect that varies with language distance. The measurement result is the more useful one: before asking whether a correspondence helps, ask how much of the score would survive without it.
Chat is not available.
Successful Page Load