When Is a Verification Layer Worth Its Cost? Operating Requirements of Certify-or-Decline Verification for Open-Weight Reasoning Agents
Abhinav Gupta
Abstract
Agentic pipelines increasingly attach verification layers to decide which of their own outputs to trust, but when such a layer pays for itself, what it requires from the underlying models, and how it fails are rarely measured. We report a measurement campaign, with decision rules preregistered for three of four arms, of a step-wise certify-or-decline verifier, a multi-role agentic pipeline in which a solver's answer is rewritten into typed proof steps and audited by specialized judges, run entirely on open-weight models (Qwen3 8B/14B/32B, gpt-oss-20b, Phi-4) across GPQA, AIME-2025, and MATH-500. Four findings decide when the layer is worth attaching. First, verification pays in proportion to how unreliable the base model is: certify-or-decline lifts precision by 31 points over answering everything on GPQA and by 46 on AIME-2025, but buys at most 5 points on saturated MATH-500 while discarding 13 to 27\% of the problems. Second, at a compute-matched budget a self-consistency selector reaches comparable precision at the verifier's operating point; the verifier leads on every seed but never significantly, and modest extra sampling budget closes the gap, so the layer's value lies in its fine-grained confidence signal and low-coverage regime rather than in dominating cheap sampling. Third, the dominant failure mode is silent: on small models most problems die at proof construction before any judge runs, and the pipeline reports these as principled declines, so a deployer would see calibrated abstention where the verifier was in fact never consulted. Fourth, certification is weaker than it looks. Audited by strong external graders for whether the accepted proof chain establishes its conclusion, proof-soundness runs 20 to 31 points below key-match precision, and open-weight judges of the certifying class cannot perform that audit at all, agreeing at chance ($\kappa \approx 0$) from 14B through 32B. We state operating requirements implied by these results and release run artifacts, scripts, and per-arm data covering every figure in this paper.
Chat is not available.
Successful Page Load