The Narrative Gap: Can LLMs help us Navigate Diverse Narratives Across Languages?
Abstract
Large language models increasingly mediate how individuals seek information about complex, often contested topics. However, no existing benchmark evaluates whether models can navigate the diverse narratives that develop around the same event across languages and regions. We argue that to support diverse information seeking, the system requires four core information navigation capabilities: differentiating narratives, presenting diverse views, resisting confirmatory queries, and achieving linguistic information parity. We introduce \textsc{NarrativeBench}, a benchmark designed to evaluate these capabilities jointly in real-world settings. \textsc{NarrativeBench} comprises 588 news articles in 17 languages from 26 regions spanning 12 geopolitical conflicts. To evaluate the models we curated 213 grounded queries from Reddit resulting in 1917 queries across 9 languages. The unit of evaluation is the model's ability to navigate diversity, not the truth of any single account. Evaluating 14 models, we find current multilingual LLMs (1) are unable to navigate diverse multilingual narratives, (2) lack the ability present diverse narratives across multilingual contexts, (3) have asymmetry in performance across different languages, (4) reduce diversity when faced with confirmatory queries which may exacerbate linguistic filter bubbles, echo chambers and reduce common ground across regions. \textsc{NarrativeBench} provides a capability-grounded foundation for building information systems that can promote equitable information access, reduce polarization, and enhance democratic discourse.