When More Runs Cannot Help: Fixed-Budget Stopping for Agent Evaluation
Abstract
Agent regression evaluation repeats a test because separate runs can produce different categorical decisions. How many runs are enough, and when can more runs no longer change the answer? We propose a fixed-budget procedure. It takes one declared categorical decision, such as an agent route or judge verdict, from each run, compares runs in disjoint pairs, and uses a Wilson interval to place the disagreement rate below a declared tolerance, above it, or leave it undecided. We then derive the largest number of disagreements a budget can absorb and still qualify. At a 5% tolerance and 146 pairs that number is two, so a third disagreement makes qualification impossible and collection can stop. In a 40-condition maintenance evaluation two conditions crossed it, and stopping there would have avoided 233 pairs: 79.8% of those conditions' budget and 4.0% of the whole workload. More runs can also answer the wrong question. Extra pairs drawn during one evaluation period pin down that period's disagreement rate rather than the average across independently repeated periods, so when that average sits at the tolerance the chance of a wrong conclusion about whether the average is below or above tolerance approaches one as the budget grows. The dependence settings are stress tests, not measurements. In practice, stop once disagreements exceed what the budget can absorb, and repeat collection in independent periods when the claim reaches across time or provider conditions.