The Grader Is Part of the Experiment: Canary-Gated Acceptance for Agent-Written Formal Proofs
Abstract
Agent evaluations for formal mathematics are graded by software, and that software can be wrong in either direction: a grader bug can reject valid proofs and deflate a score, or accept invalid ones and inflate it. The artifact defended here is an acceptance pipeline for agent-generated Lean 4 proofs in which the grader itself is under test. Four guards gate every candidate: a column-precise patch that must compile, an axiom guard that checks the kernel's axiom report for the edited declaration, a statement guard that compares the patched declaration against a snapshot of the original statement, and a reverse-dependency build of the modules that import the target. Three canary controls run alongside every evaluation: a verbatim library proof that must be accepted, a sorry submission that must be rejected by the axiom guard, and a weakened statement that must be rejected by the statement guard. During the pipeline's first deployment on a 24-task pilot over a sorry-free open-source Lean 4 library, the control matrix and its instrument checks caught five real grader bugs; four depressed scores through false rejection or degraded candidates, and one made the statement guard blind to statement weakening, an inflation defect exposed by the negative control. After the fixes, the pilot measured pass@1 of 1/24 blind and 2/24 goal-conditioned for a frontier model, one shot, no retries, and both numbers reproduced exactly under a deterministic replay and again under a replay against the original library tree, which falsified a contamination hypothesis about sibling proof holes. The contribution is the canary-gated acceptance discipline and its receipts, not a claim about model capability.