Quality-Diversity Search for Coverage Gaps in Long-Horizon Agent Safety Evaluation
Abstract
Fixed safety suites test known scenarios, and unconstrained red teaming returns many versions of the same failure. Long-horizon agents add a third problem: a small change to an early decision can produce a failure several tool calls later that no single prompt describes. We define a \emph{coverage gap} as a safety-relevant behavioral cell that a bounded, valid interaction policy reaches reproducibly and that the fixed suite and random trials do not. We propose a quality-diversity archive over agent trajectories. The genome is an intervention policy applied at a saved state, the phenotype is a trajectory descriptor, and each cell keeps a Pareto set over failure severity, reproducibility, and search cost. Counterfactual state forks are the mutation operators, and descriptor refinement splits cells that contain distinct mechanisms. Prior quality-diversity red teaming evolves single prompts; prior trajectory replay attributes failures in single runs. The archive combines the two into a search for reproducible long-horizon failures and the paths that reach them.