DIVER: Diversity-aware Agent Evaluation for Reliable Failure Discovery
Abstract
Fixed evaluation suites are effective at detecting regressions but weak instruments for discovery: they exercise only the scenarios specified in advance, and they detect only the failures their graders were designed to catch. Adaptive evaluation addresses that gap, using earlier executions to decide what to run next. When automated graders steer that decision, however, the search tends to concentrate on a defect it has already found, on parts of the scenario space the grader systematically misreads, or on scenarios no real user would issue. Flag counts then rise while the number of distinct, fixable defects does not. We propose DIVER, a training-free selection policy that discounts flags not corroborated by a second detector family, favours parts of the scenario space whose failures differ from one another, and withdraws budget from a part of the space once its findings begin to repeat. We evaluate four selection policies against an unmodified production document-generation agent: uniform random sampling, two flag-chasing baselines that maximize the detector's flag rate, and DIVER. The agent is driven through its real orchestration and tool stack rather than a mock, over 960 executions spanning grounded and ungrounded, single-turn and multi-turn scenarios in which a simulated user issues follow-ups and corrections. The generator, simulator, target agent, detector and record store stay fixed so that only the selection rule varies, and every flagged candidate goes to independent review. Both flag-chasing policies flag more executions than uniform random sampling yet return fewer distinct defects. DIVER returns the most distinct defects per flagged candidate reviewed, which is 16% more than uniform sampling and 96% more than the LLM steerer, at the lowest cost per confirmed defect.