The Ruler and the Judge: Benchmark-Conditional Evaluation of LLM-as-a-Judge
Abstract
LLM judges are increasingly used to rank systems and validate claimed improvements, but their validation often reduces to a single question: how well do judge scores correlate with aggregated human ratings? We argue that this question is too fragile. A judge can correlate with human averages while collapsing the rating scale, compressing meaningful margins, or confidently separating systems that humans do not distinguish. Furthermore, a correlation-based workflow always returns a leaderboard, even when the human benchmark cannot support one. We propose a benchmark-conditional measurement framework for LLM-judge validation. Stage 1 audits the human benchmark as a reference ruler using a many-facet ordered Rasch model, estimating rater/item fit, conditional precision, and supported pairwise resolution. Stage 2 calibrates each judge as a linked measurement instrument with judge-specific ordinal thresholds. Stage 3 evaluates out-of-sample system-level fidelity using normalized signed-margin fidelity, false-separation diagnostics, empirical transfer intervals, and distributional EMD. On SummEval and HiTZ-BASSE, the framework yields three outcomes: selection, warning, and refusal, depending on the ruler and judge pool. Calibrated predictions improve observable human-score distribution EMD in 57/81 judge–dimension cases, while the remaining failures are concentrated in the regions flagged by the ruler or judge-instrument audits. Together, these results instantiate benchmark-conditional LLM-judge validation as a protocol for deciding when judge-based claims are warranted, uncertain, or unsupported, addressing a central blind spot of correlation-based judge evaluation.