Diversity You Cannot See From the Mean: Calibrated Detection of Latent Failure-Mode Structure in Open-Model Ecosystems
Abstract
Mean-level diversity statistics for a panel of models, such as double-fault and pairwise disagreement, score behavioural spread from how often members fail, or answer, alike. We show that for failure outcomes this class of statistic is degenerate: mean pairwise co-failure, and with it mean pairwise disagreement, is a deterministic function of item difficulty alone and carries no information about shared behaviour once difficulty is accounted for. On 1,228--1,362 open language models across five benchmarks, the "excess" co-failure a naive computation reports is identically zero under the correct null. Calibrating against that null, the exact fit-free conditional distribution implied by Rasch sufficiency, reveals structure invisible to the mean in both directions: the raw participation ratio of 1.5--3.6, which reads as near-total redundancy, becomes 17--119 once calibrated, while a residual correlation of 2.9--11.9x its null remains and is concentrated in 4--12 latent behavioural dimensions rather than one shared factor. We then show the degeneracy is not merely an identity but reverses a decision: on a 2x2 grid of synthetic red-teaming suites with known ground truth, double-fault ranks a suite with planted shared failure modes as more diverse than an independent one in both difficulty regimes, and pairwise disagreement does so whenever probe potency is heterogeneous, while the calibrated statistic is correct in every cell. We describe the calibration as a drop-in diagnostic for diversity-aware benchmarks rather than a replacement for search-based diversity methods.