LocalTurn: Native Multi-Turn Safety Evaluation in Emirati Arabic, Bangla, Sinhala, and Yoruba
Abdullah Ibne Hanif Arean ⋅ Fatma Egal ⋅ Fathima Raihaan Ihsan ⋅ Victor T Olufemi ⋅ Naiyarah Hussain
Abstract
Frontier LLM safety is reported as an aggregate: a harmless-response rate, sometimes broken out by language, usually obtained by translating an English prompt set. This treats language as a label rather than as a variable. LocalTurn comprises 256 three-turn crescendo interactions authored by native speakers in Emirati Arabic (51), Bangla (89), Sinhala (50), and Yoruba (66), with 768 blind annotations from 31 native annotators and 6,144 matched ratings of 2,048 trajectories from 30 native evaluators on a 0--5 severity scale, eight commercial systems accessed through their deployed consumer interfaces. Three findings show what an aggregate hides. First, harm category composition differed significantly across languages (Cram\'er's $V = 0.360$) and rankings reorder by language and by harm category, so pooled language columns do not measure the same construct. Second, the same contributor population agreed moderately on the harm of an \emph{observed} response (Krippendorff's $\alpha = 0.470$) but near-chance on the severity of an \emph{intended} harm in a prompt ($\alpha = 0.083$), and prompt severity did not predict model failure ($|\rho| \leq 0.057$): realised output is a far more stable measurement target than anticipated intent. Third, a severity score of zero records the absence of scored harm, not a refusal. Mean severity ranged from 0.95 (Claude Sonnet 5) to 2.63 (GLM-5.2) and high-severity rates from 7.9\% to 38.8\%; four systems formed a lower-risk band separated from the other four by all 16 Holm-corrected cross-band contrasts and robust to median-of-three and single evaluator recomputation, though no contrast within the higher-risk band is supported and Claude and GPT-5.5 do not differ. Safety for these languages should be reported per language variety and per domain, with uncertainty and with repeated human judgments retained. The prompts, annotations, ratings, and trajectories will be released upon publication.
Chat is not available.
Successful Page Load