The Tokenization Tax: How Vietnamese Typology Exposes Vocabulary Design Failures in Open-Source LLMs
Abstract
Tokenizer vocabulary design is a linguistic decision with measurable production consequences. We study the tokenization tax---excess tokens consumed for non-English text relative to an English baseline---for Vietnamese, a tonal, isolating language with a diacritical Latin script spoken by over 97 million people. Across six open-source 7-9B parameter models, fertility ratios range from 0.83x (VinaLLaMA, Vietnamese-native vocabulary) to 1.93x (Mistral 7B, English-centric BPE), driven almost entirely by diacritic decomposition. A vocabulary size of ~150K multilingual tokens is sufficient for near-English Vietnamese fertility; vocabulary composition matters as much as size---a 65K English-centric vocabulary produces worse Vietnamese fertility than a 32K one. A production study on managed cloud endpoints reveals the tax is tokenizer-generation-specific: Mistral's current cloud generation uses the Tekken tokenizer (~131K multilingual vocabulary) and tokenizes Vietnamese at 1.40 tokens/word, indistinguishable from a GLM-family endpoint (1.43)---far from Mistral 7B's 2.84. Observed context-utilization differences in cloud serving are driven by window size, not fertility. All data released on Zenodo (CC-BY 4.0).