Morphology as a Constraint on Learned Tokenization: NedoTokenizer for Turkish Foundation Models
Abstract
Tokenization is a representational choice: it determines which surface patterns share parameters and how much source text fits into a fixed model budget. For morphologically rich languages, however, statistical compression and linguistic structure need not induce the same boundaries. We present NedoTokenizer, a Turkish tokenizer that treats morphology as a constraint on a learned surface representation rather than as the token inventory itself. A symbolic analyzer with sentence-level ambiguity resolution proposes internal word boundaries; a learned 32K exact-string vocabulary may split further inside those regions but cannot cross selected morphological cuts, while a complete byte fallback preserves arbitrary input. We evaluate the linguistic analysis and the actual language-model stream separately. On 210,960 valid TrMorphTester forms, analyzer cuts reach 0.9606 boundary micro-F1, while emitted IDs reach 0.8215 because the learned vocabulary introduces additional within-morpheme splits. In 36 H100 language-model runs over six native source-exact representations, byte and character models obtain lower Turkish bits-per-byte (BPB) under source-paired equal data, whereas NedoTokenizer obtains the lowest BPB under equal compute. A 2×2 H100 factorial shows that the runtime morphology constraint improves BPB in every matched comparison, while morphology-informed vocabulary induction is budget-dependent. A separate matched MorphBPE-style control improves both boundary F1 and BPB over ordinary BPE. Native Turkish Universal Dependencies experiments show competitive, but not universally best, morphosyntactic transfer. The results support a narrow principle: linguistic structure can productively constrain learned token representations without requiring one token per morpheme.