Replay Is Not Robustness: Composition Stress Tests for LLM-Judge Evaluation
Khaled AlNuaimi ⋅ Andreas Henschel ⋅ Gautier Marti
Abstract
Diversity-driven robustness search increasingly relies on LLM judges, yet their agreement scores summarize a particular item set and panel. We formulate item and panel composition as a structured robustness space whose axes operationalize item and evaluator diversity, and present a counterfactual stress-testing protocol rather than a new agreement coefficient. The protocol searches recorded and controlled recompositions for failures and maps each diagnostic to a stop rule. Across four regulated-disclosure corpora containing 8,607 analyzed items, a retained 199-item record from one corpus gives Fleiss $\kappa=0.6126$, versus 0.4566 on all 2,154 items; mean pairwise raw agreement similarly rises from 0.6990 to 0.7889. Exact replay recovers every identifier, but only 4/100,000 design-matched draws reach the recorded score, showing that replay does not establish typicality. Panel perturbations expose a complementary failure: for mean pairwise Cohen $\kappa$, a pair's mutual edge cancels, so their order depends on agreement with the remaining members. In a recorded four-model panel, 2/4 single-member omissions reverse a retained order; each reversal persists in at least 99.9% of meeting resamples. Adding a recorded adjacent-version entrant shifts one incumbent gap by +0.057 (CI [0.0469,0.0686]), although only 52.75% of resamples cross in rank. A known-truth model further identifies the dependence frontier beyond which less accurate judges that err together outrank a more accurate judge. These tests help prevent robustness claims from becoming artifacts of evaluator composition, but characterize detectable failures rather than their prevalence; agreement remains distinct from accuracy without reference labels or an explicit, testable error model.
Chat is not available.
Successful Page Load