Failure Atlases: Quality-Diversity Illumination of Interpretable Subpopulation Failures under Distribution Shift
Abstract
Distribution shift can degrade a classifier unevenly, concentrating new errors in specific parts of the target population even when aggregate performance changes only modestly. Auditing such failures requires recovering the structure of this heterogeneity, yet existing subgroup discovery methods typically return a worst-case subgroup or a ranked collection of slices without describing how failures vary across the population. We formulate distribution-shift auditing as a quality-diversity illumination problem and use MAP-Elites to construct interpretable failure atlases. Candidate subpopulations are conjunctions of feature predicates, scored by target-domain error under a minimum-support constraint and organized according to their size and demographic distinctiveness. Across 6 distribution-shift settings and 3 model families, MAP-Elites produces higher held-out quality-diversity scores and more statistically validated failure regions than random rule sampling, worst-case hill climbing, and greedy beam subgroup discovery. A cover-based diverse subgroup discovery baseline performs considerably closer to MAP-Elites, indicating that explicit diversity pressure accounts for much of the improvement, with MAP-Elites retaining the advantage of organizing failures directly over prespecified audit dimensions. We further find that evaluating discovered subgroups on the same sample used for search can substantially inflate quality-diversity scores, particularly at low minimum-support thresholds, motivating held-out evaluation of the resulting archive. These results show that quality-diversity search can turn subgroup failure discovery into a structured map of where model errors concentrate under distribution shift, giving auditors a broader view of failure severity, coverage, and heterogeneity than a ranked list of high-error groups.