Romanization Flattens the Field: Script and Metric Sensitivity in Sinhala SLM Evaluation
Abstract
Language model benchmarks for Sinhala are written in the formal script, while the people who would use a model mostly type an ad hoc Romanized register in Latin letters. We study that mismatch for Sinhala, an Indo-Aryan language with over 17 million speakers. A recent benchmark reported models are about 312 times worse on Romanized than Sinhala script text, and no effect of scale. Remeasuring 31 checkpoints, we reproduce those perplexities and find almost all of the ratio is the unit of account. Word counts are identical within a parallel pair, and in matched units the median checkpoint assigns only 2.3% more loss to the same content, while 12 of 24 assign less. Measured tokenizer independently, scale does predict Sinhala script ability. In the first downstream evaluation of Sinhala script variation, 9,479 items and eleven instruction tuned checkpoints over 208,538 generations with the romanized side produced by a deterministic rule based transliterator, the effect is not a uniform drop but a flattening. A 22.4-point spread in Sinhala script accuracy becomes 6.7, and Romanized accuracy tracks it at a slope of 0.45 across benchmark cells and 0.22 across checkpoints. On offensive language detection, accuracy barely moves while Matthews correlation loses over half its value. We close with five methodological practices for evaluation under script variation.