The Silent Regression Problem in Agent Evaluation
Abstract
When a developer edits an agent by tweaking a prompt, swapping a tool, or changing a reasoning strategy, the core question is simple: did the agent actually improve? In practice, developers often answer this by comparing noisy success rates on a small set of tasks. As a result, real performance degradations are statistically indistinguishable from noise and slip through silently. We reframe development time verification as a paired change detection problem and make it measurable. By injecting agent edits of a known sign and magnitude into a tool agent benchmark equipped with a programmatic checker, we score how reliably different verifiers and decision rules recover the true outcome under a strict labeling budget. We introduce a controlled regression benchmark, a reliable evaluation protocol, and two diagnostic metrics. Our protocol combines paired evaluation, a calibrated verifier that can abstain, and anytime valid sequential testing. Our diagnostic metrics, the silent regression rate and the false accept gap, quantify the difference between using transcript reading LLM judges versus environment grounded checks. In a controlled study with known ground truth, we find that pairing roughly halves the silent regression rate compared to a standard unpaired comparison. Additionally, using calibrated abstention recovers most of the grounded checker's detection capability at a fraction of its query cost.