Romanized Hindi Reveals a Model-Dependent Safety Asymmetry in Open-Weight Language Models
Abstract
Romanized Hindi—Hindi typed in Latin script, often mixed with English—is an important condition for language-model safety evaluation. We ask how responses to harmful requests differ from their English counterparts. We build a shared, frozen bank of 504 meaning-matched English / Romanized-Hindi prompt pairs over four harm categories and three adversarial framings, certified before any target model is contacted, and evaluate three open-weight targets on the identical bank. Scoring responses on a 0–3 information-assistance scale, we find a large but strongly model-dependent asymmetry. Qwen3-30B-A3B is scored as withholding assistance on 65.9% of English prompts but only 22.8% of their Romanized-Hindi counterparts (+43.1 pp, 95% CI [38.5, 47.6]), with 85 of 504 pairs moving from scores 0–1 in English to 3 in Romanized Hindi; the same bank yields +4.6 pp for GPT-OSS-20B and −1.8 pp for Nemotron-3-Nano. The Qwen gap remains large under a second judge, and an exploratory single-co-author audit of 354 verified items yields the same qualitative ordering of model gaps. Because both conditions use Latin script, the result is a language-conditioned contrast rather than a script effect; it does not identify a mechanism. English-only measurements may therefore fail to characterize Romanized-Hindi behavior.