Small but Mighty: Test Suites for Self-Evolving Agents via Quality-Diversity Evolutionary Search
Abstract
A self-evolving agent rewrites its own harness and must decide which rewrites to keep. A rewrite that silently changes behavior is inherited by every later episode, while a verification signal that rejects behavior-preserving rewrites stalls the evolution it was meant to protect. Rollout feedback spends an agent episode per decision, and the repository's own tests, the obvious free signal, reject only a fifth of the rewrites that do change behavior. We generate the signal instead, and treat generating it as a quality–diversity problem: suite size becomes a behavior descriptor rather than a penalty inside the objective, so quality is pursued at every size and a small, strong suite is reachable rather than traded for. Reaching a given size means instructing the LLM to write that many tests, and an LLM follows a count instruction with a systematic bias, so the size requested and the size written are not the same number. TeSearch learns that bias online as a transition kernel over sizes and corrects for it, so samples land where the search intended and the bandit that spends the budget scores what actually happened. Across three backbones, TeSearch delivers the highest-quality suites of any method compared and the smallest of any LLM-based one. Used as the verification signal of a self-evolving agent, it rejects 57.6% of behavior-changing rewrites against 22.2% for the repository's tests, at a 3.8% false-alarm rate and 3.5 s per decision.