When Expert Disagreement Hurts: Auditing Prestige-Sensitive Revision in LLM Decision Pipelines
Abstract
Language models are increasingly used in decision pipelines where an initial answer is revised after another agent disagrees. In such settings, revision can reflect useful reconsideration, ordinary prompt instability, or sensitivity to the perceived prestige of the disagreeing source. We introduce a controlled two-pass audit that separates these effects by holding the question and evidence fixed while varying only the revision context: neutral reflection, anonymous disagreement, and expert-labeled disagreement. We evaluate five instruction-tuned models on 1,989 evidence-grounded binary decisions from medicine, scientific claim verification, and contract reasoning. Disagreement consistently induces reversals beyond drift, while the incremental effect of an expert label varies across models. More importantly, these reversals are often harmful: under expert-labeled disagreement, 70.1% of reversals aggregated across models move from correct to wrong answers. In a branch-paired API subset, low-prestige disagreement is weaker than expert disagreement, and a simple evidence-gated revision prompt reduces harmful reversals while improving final accuracy. For open-weight models, teacher-forced Yes/No scores reveal that many output reversals are not accompanied by corresponding preference-sign changes. These results make prestige-sensitive revision a concrete, reproducible, and actionable target for LLM evaluation.