Weight-Space Signatures of Tokenization Cost and Morphology in Small Language Models
Abstract
The tokenizer is the first and least-examined design decision that fixes the linguistic medium a foundation model perceives. This paper audits that decision from three angles across a zoo of 19 open small language models spanning 9 tokenizer lineages. On 38 languages of parallel content (Tatoeba), the known tokenization-premium effect replicates and extends, costing a zoo-median 1.86 times more tokens in a non-English language than in English, rising to 7.1 times for Tamil and Telugu. Whether this upstream inequality is visible in the trained weights themselves, via embedding-space anisotropy concentrating on non-Latin-script tokens, is tested next. The effect is present but small and not statistically distinguishable from zero at the population level (12 of 19 models, bootstrap 95% CI [-0.0055, +0.0445]), while the two massively multilingual models in the zoo (BLOOM, BLOOMZ) show a clear reversal (-0.077), and the two axes are nearly uncorrelated at the language level (r = +0.10 across 36 languages). Fertility is where the measurable, robust inequality lives, and embedding anisotropy is a smaller, training-mixture-contingent effect. Finally, English inflectional families from UniMorph (e.g. run, runs, running, ran) show substantially higher within-family cosine similarity than random word pairs across a 6-model subset (mean gap +0.405, positive in 6 of 6 models, exact binomial p = 0.031), surviving a frequency-matched control and extending, on a small sample, to German and Spanish. Weight geometry does encode compositional and morphological relatedness, even though it does not clearly encode script-level training coverage, so embedding geometry is selectively, not uniformly, informative about linguistic structure.