When Failure Descriptors Hide Repair Differences: A Held-Out Audit for Quality-Diversity Failure Archives
Abstract
Quality-diversity search organizes robustness failures with a behavior descriptor, and the number of occupied cells is read as a measure of diversity. Do failures that respond alike to individual repairs also respond alike when repairs are combined? We test this by grouping digit-classification errors according to their repair responses and reserving other repair combinations for an audit. With oracle removal of known corruptions, 47.3% of grouped logistic-regression failure pairs and 45.6% of grouped neural-classifier pairs disagree on at least one audit repair. A follow-up that applies inverse transformations without clean-image or corruption-identity access finds disagreement of 62.5% and 52.0%. Adding selected repair probes separates more of these pairs than adding an average panel of the same size, at an extra discovery cost. The decision consequence is less clear-cut: choosing one repair per group leaves an empirical success gap, but known corruption identities reduce that gap more than the full added probe panel. Whether the audit improves an evolving failure archive remains untested; the experiments show why a descriptor should be checked against the repair decisions it is intended to support before an archive relies on it.