VerifierContractBench: Does Interactive-Agent Grader Optimization Preserve Safe Decisions?
Abstract
Trajectory-level graders make interactive-agent evaluation scalable, but the graders themselves are increasingly optimized by coding agents. An improved aggregate score does not reveal whether the new grader has begun approving failures that its predecessor correctly rejected. We introduce VerifierContractBench, a benchmark for evaluating grader-optimization processes. Each episode starts from the same frozen multimodal verifier, exposes 102 human-labeled browser-agent trajectories, permits six edit–evaluate rounds under a fixed acceptance gate, and freezes the final incumbent. We compare utility-only, aggregate-safety, and instance-preserving gates across five coding meta-agent endpoints and text-only versus multimodal evidence, yielding 30 episodes. Final verifiers are evaluated on 492 task-disjoint protected trajectories. Our primary metric is a paired transition rate: among human-unsafe trajectories rejected by the starting verifier, how many are newly certified after optimization? All 26 genuinely edited verifiers introduce at least one such regression while also repairing old errors. Instance preservation lowers mean regression rate from 8.19% to 5.19%, but does not eliminate it. The benchmark contributes a rigorous protocol for grader design that reports decision transitions, rollback dependence, proposal behavior, and protected generalization rather than only final accuracy.