Güleç: Morphology-Aware Word-Level Compilation for Efficient Turkish Language Modeling
Abstract
Subword tokenization provides a general solution to open-vocabulary language modeling, but in morphologically rich languages it can substantially increase Transformer sequence length by representing a single orthographic word with multiple subword positions. We introduce Güleç, a morphology-aware language model architecture for Turkish that compiles lexical fragments and morphological structure into a single Transformer position per word. Güleç combines a lexical composer, a morphology encoder, and a learned low-rank morphological transformation to construct one word-level representation, while a lightweight local structured decoder preserves open-vocabulary generation outside the global Transformer sequence. We evaluate Güleç against parameter-matched standard BPE Transformers through mechanistic ablations, controlled common-head experiments, full surface-level evaluation, and systems benchmarks. At approximately 300M parameters, a common-head comparison isolates the input representation and yields a validation cross-entropy of 4.783 for Güleç versus 4.991 for the BPE baseline. In the full model, under an explicitly restricted train-only empirical realizer (R_{\mathrm{train}}), Güleç achieves 7.717 surface NLL/word compared with 11.740 for the standard Transformer on the same 431,900 validation surface events, a 34.27% reduction. In the corresponding all-target training run, Güleç processes 49.55k input words/s versus 44.18k words/s and uses 11.83 GiB versus 19.67 GiB peak Torch memory. A batch-1 KV-cache benchmark additionally shows a 1.34× decode-throughput advantage in surface words per second under the tested implementation. These results support a simple architectural hypothesis: for Turkish, morphology need not consume additional global Transformer positions. Instead, lexical and morphological structure can be compiled locally into word-level states, improving predictive quality while reducing the sequence and memory costs associated with subword modeling.