Four Silent Failure Modes in Measuring Multi-Agent LLM Verification
Abstract
Multi-agent LLM systems increasingly rely on verification-one model checking another's output, or consensus across instances to catch errors. We ask a prior question: can we even measure how well verification works? Testing three verifiers (GPT-40, Claude Sonnet 4.6, Llama-3.3-70B) on whether they accept incorrect peer answers, we find the measurement itself is riddled with silent failure modes. We document four, each invisible in aggregate metrics and each surfaced only by logging and auditing raw responses: (i) forcing prompts that safety-trained models refuse, silently replacing wrong answers with correct ones; (ii) invalid error sets, where 66% of auto-harvested "errors" were not errors; (iii) token-budget truncation that cuts deliberative verifiers off mid-reasoning, shifting their scores by up to 80% relative; and (iv) gold-contaminated grading, where showing the grader the answer made it judge correctness instead of transcribing behavior. After correcting all four, a pre-registered protocol on the human-verified SimpleQA Verified benchmark shows that verification which catches ~99% of errors on easy queries collapses on rare ones (acceptance 0.30-0.81, all p < 10 ^ - 4 ) with a 51-percentage-point spread between verifiers. Peer-resistance is not a fixed model property; it is bounded by shared knowledge, and its measurement is fragile by default.