When the Verifier Is Wrong: Red-Teaming Agent Verification Where Ground Truth Exists
Abstract
Reliable verification requires a trustworthy connection between the verification signal and the outcome being evaluated. Open-ended agentic tasks rarely provide complete ground truth, which makes verifier failures difficult to detect. We use a closed-loop protein-design agent as a controlled testbed because nearly every available action has a recorded experimental outcome. This outcome table lets us evaluate the verifier itself against ground truth. Our trust gate certifies a scorer that has memorized the observations used for evaluation. The memorized scorer ranks above an honest scorer and passes the gate in every trial, yet performs no better than random and loses 1.05 good candidates per batch. Forward validation prevents this leakage attack without inspecting the scorer or reducing the honest scorer's pass rate. We then construct a pool-side attacker that leaves every measured candidate unchanged and corrupts only the unmeasured pool. Its gate statistics match the honest scorer to three decimal places, yet it loses 1.24 good candidates per batch. Sequentially measuring top-ranked candidates detects 99.2\% of the pool-side attacks, but detection alone recovers only 10\% of the lost value. Replacing random fallback with a model trained on the agent's own measurements raises recovery to 28\%. As the campaign grows, that in-house model eventually outperforms the supplied scorer and removes the attack surface. Across 41 tasks, we identify four additional verifier pathologies and withdraw a fifth that disappears after expanding the evaluation set. Complete-outcome domains provide a practical testbed for agent verification because they expose failures in both the agent and the verifier.