Authority, Truth, and Citation Bias: A Multi-Domain Benchmark for Studying Citation Deference in Large Language Models
Abstract
Large language models are increasingly deployed in settings where outputs are grounded in cited sources, yet whether citation presence itself, independent of factual content, shapes model behavior remains poorly understood. We introduce AuthorityBench, a benchmark of 220,564 prompts designed to isolate the effect of citation presence on factual accuracy in large language models. Citations appear passively in prompts, models receive no instruction to trust or defer to them, allowing us to measure whether citation presence alone cause models to give wrong answers. The benchmark independently manipulates two variables: whether the claim is true or false, and whether the accompanying citation is real or fabricated. This yields four conditions tested across four domains (general knowledge, science, law, and medicine), with controlled variation over 40 prompt templates, four venue prestige tiers, and author names from six surname regions. We define hallucination as any response contradicting the ground-truth label, wrongly denying a true claim or wrongly affirming a false one, capturing both the familiar failure of endorsing misinformation and the novel failure of rejecting correct facts under authority pressure; refused responses are excluded from this count. Across all seven models, adding any citation, real or fabricated, increases hallucination above baseline. The effect is largest when a fabricated citation accompanies a true claim: models that answer correctly without a citation are pushed toward wrong answers by the authority signal alone, with hallucination rates rising 3 to 22 percentage points and reaching 35--77\% in general knowledge. Legal claims are consistently robust; venue prestige and author demographics have negligible effects; and larger models are no more resistant than smaller ones. All code, datasets, prompt templates, and evaluation resources are available at https://anonymous.4open.science/r/AuthorityBench-3C0C.