WAHM: Measuring Hallucination Drift Across Arabic Varieties
Mujtaba farhan ⋅ Muhammad K Khan
Abstract
Arabic language models are commonly evaluated in Modern Standard Arabic (MSA), while everyday communication spans regional varieties that differ substantially in vocabulary, morphology, and syntax. We introduce WAHM, a paired benchmark for measuring how open-generation factual hallucination changes when the same information need is expressed in MSA, Gulf, Egyptian, Levantine, or Sudanese Arabic. WAHM contains 300 questions aligned across five Arabic varieties and introduces the Hallucination Drift Score (HDS), which measures the change in hallucination rate between a dialect and its matched MSA baseline. We complement HDS with hallucination-set intersection-over-union (IoU) to measure whether the same questions fail across varieties. Six Arabic-centric and multilingual models are evaluated using conservative degeneration screening followed by a question-aware AraBERT factual-hallucination judge, with blinded human validation. Results show that MSA factual reliability and dialect robustness are distinct: ALLaM has the lowest MSA hallucination rate yet exhibits a $+10.1$ percentage-point drift on Sudanese, while Jais shows comparatively small aggregate drift across dialects. Error-set overlap further reveals that similar hallucination rates can conceal substantial question-level turnover. WAHM provides a controlled framework for evaluating factual reliability beyond MSA and for distinguishing changes in hallucination frequency from changes in which questions fail.
Chat is not available.
Successful Page Load