VerifierContractBench: Auditing Safety Regressions in AI-Optimized Oversight
Abstract
Automated verifiers increasingly determine whether agent behavior is rewarded, entered into trusted datasets, or approved for deployment. Coding agents can now modify these evaluators, but higher aggregate accuracy can conceal a withdrawn protection: an unsafe trajectory previously rejected may become certified. We introduce VerifierContractBench, a benchmark that audits the complete verifier-improvement episode. Each episode starts with the same frozen multimodal verifier, provides a coding meta-agent with 102 expert-labeled web-agent trajectories, permits six edit-evaluate-rollback rounds, and freezes the incumbent. Across five endpoints, text and multimodal evidence, three promotion gates, and three independent repetitions, we run 90 episodes, 540 rounds, and score 44,280 protected episode-trajectory outcomes. Our primary new unsafe-certification rate (NUCR) identifies human-unsafe trajectories that are rejected by the starting verifier but are newly certified after optimization. All 78 genuinely edited final verifiers introduce at least one such regression; none improves protected balanced accuracy without one. Removing every case unstable across five fresh byte-identical baseline audits leaves all 78 positive; two further rechecks retain a strict 3-of-3 regression in 77 of 78. Utility selection yields 8.81% NUCR, versus 5.52% for aggregate safety and 5.63% for instance preservation. Relative to utility, the reductions are 3.29 and 3.19 percentage points; hierarchical 95% intervals are [-5.67, -1.10] and [-5.44, -1.08]. The two safety gates do not differ detectably (instance minus aggregate: +0.10 points, [-2.29, +2.42]). Mean aggregate risk is similar across repetitions while exposed-case identity is not. The benchmark specifies requirements for paired transitions, protected testing, risk stratification, inference-stability controls, and rollback provenance evidence for trustworthy evaluator updates.