Nominal Coverage Is Not Failure Diversity: Human-Verified Repertoires in Adaptive Robustness Search
Abstract
Robustness evaluations often treat nominal scenario or category coverage as evidence of diverse failure discovery, yet the same covered categories can yield very different sets of verified failures. We audit this gap through a quality-diversity lens using a prospectively collected human-verification experiment. Two frozen music representations are tested over six controlled transformations by three policies under matched prospective query budgets: human-grounded \textbf{ACTIVE} prioritizes predicted human--model disagreement, \textbf{STATIC} ranks target-model separation, and \textbf{RANDOM} uses seeded priorities. A behavioral niche is defined by the target representation and the transformation it prefers; a niche is verified when a majority of later human judgments oppose that preference. Although all policies satisfy the same minimum transformation-coverage rule, their headline searched/verified niche counts are 8/7 for ACTIVE, 5/5 for STATIC, and 9/4 for RANDOM; mean disagreement support is 0.875, 0.600, and 0.500. Under the deployed joint allocation, RANDOM accumulates the broadest two-round verified archive (8 niches) despite weaker headline yield, while matched replay shows that adaptive state and score changes need not change the repertoire searched next. The analysis separates nominal coverage, searched breadth, failure yield, current verified repertoire, cumulative archive, and update-to-repertoire change over selections and outcomes frozen prospectively.