Position: Morphology, Not Frequency, Should Govern Vocabulary Allocation
Abstract
Recent works have established that tokenization imposes a structural token tax on most of the world's languages. Identical content costs several times more tokens outside English, and token counts track downstream accuracy. Those studies repeatedly observe that high-fertility tokenizers fragment morphology, yet they route around it. The remedies they propose are larger vocabularies and broader script coverage. Where morphologically aware tokenization is named at all, it is named in a closing clause and barely defended. We argue that the neglected recommendation is the right one, though not in the form it is usually stated. Frequency-driven vocabulary construction encodes an implicit linguistic theory, namely that statistically frequent strings are the units of meaning. That theory degrades as a language's morphological synthesis increases. We test this using 12 tokenizers over 21 languages, gold morpheme segmentations for 19 of them, a fixed-budget inference study over two open-weight models, and a matched-budget allocator prototype. Four findings shape our position. First, fertility is largely a budget fact. The vocabulary a tokenizer spends on a script explains most of the variance in fertility, and boundary alignment adds nothing once that budget is controlled. Second, alignment is nevertheless a distinct axis that budget does not buy, and the way it relates to fragmentation is conditioned on typology. Third, the budgeted evaluation reproduces the link between fertility and accuracy on Indic languages, while returning a null for alignment beyond fertility. Fourth, a prototype allocator recovers 24-97\% more boundary alignment at a matched vocabulary budget for a 4-12\% fertility cost. We conclude that tokenizers should be evaluated on morphological alignment and cross-typological equity rather than on compression alone.