Ambiguity-Aware Evaluation for Automated Fact-Checking
Abstract
Automated fact-checking systems are evaluated almost entirely by whether their predicted verdict matches a single gold label, an approach that assumes every claim has one stable interpretation and that the available evidence is always sufficient to resolve it. In practice, claims are frequently underspecified in time, place, or meaning, and retrieved evidence is often incomplete or contradictory, so this approach cannot distinguish a system that reasoned well on a genuinely ambiguous input from one that reasoned poorly but matched the label anyway. We introduce a framework that separates ambiguity in the claim and evidence from failures in the system's own reasoning and grounding, and conditions accuracy and response appropriateness on this ambiguity profile. Applied to a retrieval-based fact-checking pipeline across five benchmarks, plus a human pilot study, we find that ambiguity is pervasive and dominated by missing timeframes, that most correct verdicts on several benchmarks are backed by reasoning that would not hold up under scrutiny, and that systems grow more confident, not less, when their evidence conflicts.