What LLMs Know About Kyrgyz and What They Cannot Do: A Native-Authored Benchmark Separating Knowledge from Grammar
Abstract
Kyrgyz, a Turkic language with roughly five million speakers, is nearly absent from large language model evaluation, and where low-resource Turkic languages are measured at all the evidence is usually a translation score on a machine-translated test set, which conflates the model with the translation system used to build the test. Using a native-speaker-authored benchmark of 64 four-option items across eight linguistic phenomena, I find that frontier models share a systematic failure on the morphology of this low-resource language: they answer factual and lexical questions at or near ceiling (idioms, translation, and lexical semantics reach 100%) while dropping to 62% on the items that require producing correct inflection, with an identical score for the small and the large model in the family. This knowledge–grammar dissociation is invisible to a single aggregate accuracy, which averages the grammatical failures away; isolating and scoring morphological accuracy on its own is what exposes it, and a 300-item, corpus-grounded morphology benchmark confirms the effect at scale (gpt-4o 82.7%, gpt-4o-mini 65.7%). I then show the gap is learnable rather than intrinsic: a rule engine for Kyrgyz nominal inflection, verified against the hand-authored items, generates training data on which a 0.5B open model, adapted with LoRA parameter-efficient fine-tuning, rises from 0% to 77% at producing inflected forms for held-out words. Scoring is deterministic and every result is tested against a 25% random baseline; the benchmark, the rule engine, and all code are released so that every number here is reproducible.