KamtaCorpus: Building and Analyzing an Open Text Corpus for Kamtapuri, a Low-Resource Language of the Global South
Abstract
Kamtapuri (Rajbanshi) is an Indo-Aryan language spoken by an estimated 10–15 million people across North Bengal, Assam, and Nepal, and is officially recognized in West Bengal—yet it is almost absent from language technology and is routinely misread as a dialect of Standard Bengali because the two share the Eastern Nagari script; prior work finds that frontier LLMs overwrite it with Bengali (Das, 2026). Building technology that resists this begins with data. We present KamtaCorpus, a cleaned corpus of 348,688 words (33,922 sentences) digitized from Kamtapuri literature and manually verified by native speakers, together with two analyses of it: (i) a tokenizer analysis showing that, on domain-matched literary text, Kamtapuri costs 7–14% more tokens per word than Standard Bengali on two Indic-aware tokenizers, a gap that reverses on a general-purpose one; and (ii) a finetuning baseline in which LoRA on Qwen2.5-0.5B/1.5B cuts held-out perplexity by ~41%, establishing that the corpus is readily learnable. Permission to release the corpus for research use has been secured from the source publishers, and it will be released under CC BY 4.0; the split manifest and corpus statistics are given in this paper, and the full datasheet accompanies the full version. KamtaCorpus is a data-and-analysis foundation for studying—and ultimately countering—the erasure of a Global-South language; evaluating erasure itself is future work.