VerifierContractBench: Safety Regressions in Adaptively Improved Verifiers for Web Agents
Abstract
Interactive web agents are judged by automated verifiers that assign rewards or deployment approval. Coding meta-agents can adapt these verifiers, but higher accuracy can conceal a withdrawn protection: an expert-rejected trajectory becomes certified. We introduce VerifierContractBench, a benchmark of this adaptive oversight boundary. From one frozen verifier artifact and 102 expert-labeled trajectories, each meta-agent receives six edit--evaluate--rollback rounds; the selected verifier is then audited on 492 task-disjoint trajectories. Five hosted model--provider configurations, two evidence modes, three gates, and three repetitions yield 90 episodes and 540 rounds. A trajectory is human-unsafe to certify unless jointly labeled successful and free of unintended side effects. The new unsafe-certification rate (NUCR) measures cases that were initially rejected but were newly certified after optimization. All 78 edited finals regress; 77/78 also regress on a side-effect-labeled case, and none improves protected balanced accuracy without one. Removing cases whose baseline decision changes across five fresh calls leaves all 78 positive; in 77/78, the regression recurs in the original audit and both rechecks. Aggregate- and instance-safety gates reduce NUCR by 3.29 and 3.19 percentage points relative to utility, but neither eliminates regressions. Safer computer-use oversight requires paired transitions, task-disjoint audits, and rollback provenance.