The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining
Abstract
Language models (LMs) often hallucinate by committing to substantive answers when they should abstain. Existing methods detect hallucinations or guide abstention, but leave open how models internally decide to commit or abstain. We study this decision through mechanistic analysis, framing a class of hallucination as unsupported commitment: the model commits despite exhibiting signals of unanswerability. Using causal gating, we identify a Commit-Abstain Circuit (CAC), a sparse subset of attention heads and MLP sublayers that causally contribute to this decision. Across ten LMs (3B-14B) from five families and three benchmarks, the CAC exhibits a recurring accumulate-yet-undercorrect pattern: commitment-promoting components build up commitment in earlier layers, while abstention-promoting components act later as corrective signals that are often insufficient to overturn the accumulated commitment. Building on this finding, a lightweight policy trained on CAC activations improves decision accuracy by 12.2 points over the model's intrinsic commit-abstain margin, reduces false abstentions by 2.5 times, transfers to unseen benchmarks, and extends to larger models (27B-35B). The CAC is both diagnostic, clarifying how models overcommit, and practical, enabling improved abstention decisions. Code is provided in the supplementary material.