VerifierContractBench: Safety Regressions in Meta-Agent-Optimized Verifiers
Abstract
Coding meta-agents can improve the prompts and programs used to verify other agents, but an aggregate score can conceal newly introduced unsafe approvals. We introduce VerifierContractBench, a benchmark for safety-constrained verifier evolution. Each episode starts from the same frozen multimodal verifier, exposes 102 human-labeled web-agent trajectories, permits six edit–evaluate rounds, and freezes the final incumbent. We compare utility-only, aggregate-safety, and instance-preserving acceptance gates across five coding meta-agent endpoints and text-only versus multimodal evidence, yielding 30 episodes. Final verifiers are evaluated once on 492 task-disjoint protected trajectories from AgentRewardBench. Our primary metric measures unsafe trajectories that the starting verifier rejects but that are newly certified after optimization. Of 30 episodes, four make no accepted change; all 26 genuinely edited verifiers introduce new unsafe certifications while also repairing old ones. The instance-preserving gate lowers the mean new unsafe-certification rate from 8.19% under utility-only selection to 5.19%, but regressions remain in every edited episode. The mean protected balanced accuracy increases by only 0.82 percentage points, and no episode improves it without a new unsafe certification. The benchmark directly tests whether verifiers that evolve under agentic optimization remain faithful to human-grounded safety decisions.