Fine-Tuning, Not Pretraining, Produces the Multilingual Geometric Split
Abstract
Multilingual models are often probed by measuring how far apart their language representations sit, with collapsed representations read as language-neutral abstraction and separated ones as language-specific encoding. These probes are usually applied to fine-tuned checkpoints and reported as properties of the model. We compare four multilingual models across six typologically diverse languages before and after identical English-only SQuAD fine-tuning. The encoder-only/encoder-decoder split visible afterwards is absent before: pretrained mBERT is the most language-separated of the four (0.190 mean centroid distance), above both encoder-decoder models (0.116, 0.117), while XLM-R alone begins collapsed (0.003). Fine-tuning moves the models in opposite directions by factors of 1.3 to 45, but only mBERT crosses between regimes, and its 45-fold collapse is what produces the two-family pattern. XLM-R and mBERT end at similar values from opposite starting points, which post-fine-tuning geometry alone cannot distinguish. That geometry does not track downstream behaviour: two models in a comparable regime differ by 12.3 to 15.9 F1 on zero-shot extractive QA across three seeds, a less collapsed model outperforms the most collapsed by 9.5 F1, and geometry-performance correlations are indistinguishable from zero on two tasks. Typological alignment does not follow collapse either: under exact Mantel tests, three of four models align significantly with WALS distance, led by one of the most collapsed (r = 0.701), while the most collapsed shows none, though this is also measured post-fine-tuning. Geometry measured on an adapted checkpoint describes a model together with its adaptation, and is not a substitute for downstream evaluation.