Does Diversity Transfer? Evaluating Quality-Diversity Failure Archives Across Different Representations
Abstract
Quality-diversity (QD) robustness search aims to find many distinct model failures instead of a single worst-case failure, but what satisfies a distinct failure depends on the representation used to compare failures. We test whether diversity found under one representation remains when the same failures are evaluated under other, separately defined representations. Using Qwen2.5-7B-Instruct on GSM8K, we run five independent searches for each of three fixed input-side representations, including a native transformation taxonomy, semantic embedding deltas, and lexical TF–IDF deltas. Our primary analysis compares each QD archive with equally sized random sets of wrong failures found in the same search run. Across 15 runs, the archive is 1.435×as diverse as the matched random baseline under the representation optimized by search, but only 1.013×as diverse under held-out representations. This pattern remains after matching archive size, changing audit- cell granularity, and replacing cells with nearest-neighbor distance. A same-task example and a secondary output-side analysis show that prompt diversity does not necessarily imply equally diverse failure behavior. We therefore propose a stricter evaluation procedure that evaluates the archive under separately defined representations, compares it with a same-run size-matched baseline, and verifies the result with a clustering-free distance measure.