Project Pulli: Grounding Edge-Scale Pan-Indic Transliteration with Classical Phonotactics and In-the-Wild Golden Diagnostics
Abstract
While large language models and commercial cloud APIs achieve competitive accuracy on formal dictionary transliteration, their performance degrades severely on organic, conversational typing across the Global South. In-the-wild inputs suffer from lexical translation leakage, fortis geminate collapse, and unconstrained cross-script hallucinations. This paper presents PROJECT PULLI, an ultra-compact (approx. 11.4M parameter / 11.4 MB INT8) edge-scale transliteration engine spanning 8 Indic languages (Tamil, Malayalam, Telugu, Kannada, Gujarati, Bengali, Hindi, and Marathi) trained on a 3.2M-pair Pan-Indic corpus. Grounding South Dravidian phonology in classical Tolkāppiyam phonotactics and building upon the empirical information-theoretic limits of benchmark text supervision, PULLI addresses the gap between clean benchmark corpora and noisy conversational chat across both Dravidian and Indo-Aryan scripts through a hybrid framework: pairing a Tier-1 Core Lexicon Trie with an edge-scale sequence-to-sequence Transformer regularised by language-adaptive objectives (chat augmentation and puḷḷi structural loss for Dravidian scripts; supervised contrastive spelling-invariance loss and nuqta preservation for Indo-Aryan). On a 5,000-sample macro benchmark, PULLI-A2 reaches 57.8% Exact Match against 53.9% for a replicated INDICXLIT baseline (+3.9 pp, +195 net items; exact McNemar p = 3.5e-14). On an audited 216-sample in-the-wild golden suite verified by native speakers, PULLI achieves 71.3% Exact Match (154/216 items), outperforming a commercial production cloud API (66.7% EM, 144/216 items; +4.6 pp), while delivering under 2 ms on-device latency (approx. 250x faster than cloud API roundtrips) with zero operational infrastructure cost. PULLI establishes a statistically significant lead in Malayalam (+32.5 pp, +13 net items; exact McNemar p = 0.00098) alongside verified leads in Gujarati (+25.0 pp) and Telugu (+10.0 pp), while the commercial cloud API retains an advantage on Hindi, Marathi, and Bengali (-5.0 to -15.0 pp). The gains thus concentrate in, but are not confined to, Dravidian, and the split motivates the proposed language-adaptive routing, demonstrating that typologically grounded inductive biases offer a scalable and equitable blueprint for resource-efficient NLP.