Red-Teaming the Grader: Mutation-Validated Constraint Checking for Tool-Agent Benchmarks
Abstract
A benchmark grader is a verifier: every leaderboard entry is a claim that the grader certifies. For tool-using LLM agents scored by deterministic final-state checks, an under-specified checker silently converts wrong behavior into reported success—and because false successes hide on the passing side, no amount of failure-focused auditing will surface them. We present a field-tested protocol for stress-testing the verifier, applied post-run and pre-release to an 800-scenario tool-agent benchmark in professional non-linear video editing: (i) mutation probes that attack verified reference solutions with step drops, delete-substitutions, wrong targets, value perturbations, and string-literal swaps; (ii) a null-agent guard; (iii) a fail-closed predicate registry with per-predicate normal/boundary/adversarial fixtures enforced by a meta-test; (iv) content-hashed scorer identity; and (v) provenance-verified rescoring of all 117,600 stored run records. The protocol caught real defects. Across two probe rounds it surfaced 527 unique passing mutants—corrupted variants of verified reference solutions—and hardening added 310 constraints across the lineage with zero deletions; re-executing both probes under the final scorer and corpus identity leaves 240 unique survivors, each matched to a recorded adjudication rather than an open checker gap. A parallel predicate audit catalogued 20 defect classes; the combined defect ledger includes outright inversions such as a direction field that 115 constraints declared and the scorer ignored. A sensitivity annex bounds sensitivity to a stricter preservation contract: median score shift −2.14pp with rank correlation ρ=0.9974. No single component is new; the contribution is the integrated protocol, shipped with the benchmark as a machine-checkable defect ledger and a checklist distilled from what it caught.