Understood, Unaligned, or Unrouted? Three Failures Behind One Multilingual Refusal Rate
Abstract
A single multilingual refusal rate can fall for at least three different reasons, and standard evaluation reports them as one. Safety fails at the level of the variety, not the language: speakers of non-Latin-script languages routinely type them in Latin letters—Arabizi, Roman Urdu, Latin-script Bangla—through romanisations nobody has standardised, and the romanised form travels a different path through the model than the script it replaces. One row per language cannot show it. We separate three modes: Type 0, the model never recovers the request (comprehension failure); Type I, it recovers the request but represents no harm to act on (missing alignment); and Type II, the harm is represented yet refusal does not fire (a routing failure). Under current evaluation the three are indistinguishable and call for opposite fixes, so the field cannot know whether its remedies are aimed correctly; our position is diagnostic, not causal, and holds whichever mode dominates—the point is that current evaluation cannot say which. Both cuts are cheap: a comprehension probe isolates Type 0 on inference access alone, and a harm-recognition probe—does the model judge the same item harmful when asked?—proxies Type I versus Type II, which representation-space and circuit tests (Wang et al., 2025) sharpen where weights are available. Reading 32 studies against one criterion, six measure comprehension separately from refusal and none does so across natural-language writing systems. Re-analysing the one multilingual study that reports both, comprehension failure is the larger component (GPT-4 on Arabic chatspeak: 60.96% misunderstood against 3.46% unsafe over all prompts), yet conditioning on recovered prompts still raises unsafe compliance 6.6× from Arabic script to transliteration (16.99% vs. 2.56% of recovered prompts)—a residue no refusal rate would show. We propose a minimum standard, and this worked re-analysis, that any API-only team can adopt at the cost of one extra prompt per item.