When the Verifier Becomes the Objective: A Verification Envelope for Self-Evolving Research Agents
Abstract
Self-evolving research agents do more than answer questions: they select hypotheses, modify executable artifacts, run experiments, and use prior evaluations to decide what to try next. In this regime the verifier is part of the optimization environment. A verifier that is reliable for one pre-specified experiment can become unreliable under adaptive search through metric manipulation, winner selection, holdout reuse, latent-instance leakage, mechanism substitution, or simply a test suite that never exercises the gate it appears to protect. We develop a verification envelope that separates five assurance obligations—metric-channel integrity, adaptive statistical validity, mechanism integrity, evidence lineage, and verifier integrity—rather than collapsing them into a single “verified” bit. The strongest layer structurally separates untrusted artifact generation from trusted scoring; weaker semantic protections are labeled explicitly as scan-defended. We further introduce verifier mutation testing: deliberately disabling security- and statistics-critical mechanisms and requiring the operational verification suite itself to fail. A longitudinal audit of an operational self-evolving research system found (i) a winner-selection gate that could be forced to confirm unconditionally while both its nominal test and the entire core tier remained green, (ii) 65 of 66 historical holdout evaluations reusing the same five identifiers, and (iii) nominally fresh seeds mapping onto only 13 latent instances, contaminating the holdout after selection. After repair, all 14 targeted verifier mutations are detected. The central conclusion is methodological: reliable verification of optimizing agents is a systems property, and the verifier must itself be treated as an adversarially exposed object of verification.