HingLM: Isolating the Tokenizer's Effect for Romanized Hindi–English Code-Mixed Language
Aditya A Deshmukh ⋅ Shripad Deshmukh
Abstract
Romanized Hindi–English code-mixed text (Hinglish) is a challenging setting for subword tokenization, as English-based tokenizers can break Hindi words into multiple small chunks, thereby increasing sequence length and the computational demand of processing code-mixed text. We study this effect through a controlled tokenizer comparison, following the methodology of Rust et al. (2021), in the interleaved, romanized setting, which has not been explored in previous controlled studies. A BPE tokenizer trained on Hinglish data reduces the number of tokens by 24.9% compared to GPT-2 on held-out text and allows the model to encode $\sim$33% more characters per token. This advantage increases with the level of code-mixing, from 6.2% in lightly code-mixed sentences to 29.9% in heavily code-mixed sentences. We also assess the two tokenizers in a controlled 2×2 design spanning tokenizer choice and initialization on token-level language identification. The custom tokenizer shows a 3.36-point improvement in macro-F1 with pretrained initialization across ten seeds and a 4.56-point improvement from scratch. In our experimental setting, the tokenizer effect is therefore larger than the effect of pretraining.
Chat is not available.
Successful Page Load