A Second Judge is Not a Second Opinion: Correlated Error and Silent Failure in LLM Verifiers
Varad Vishwarupe ⋅ Nigel Shadbolt ⋅ Meshari M Alwazae
Abstract
Two open-weight LLM judges labelled the same 54 utterances under byte-identical prompts. They agreed with each other at Cohen's $\kappa = 0.879$, with the human coders at $\kappa \approx 0.19$, and with two frontier judges at $\kappa \approx 0.10$. Running both and taking the majority would therefore have bought almost no independent evidence, only a second copy of the same error. We report this from a controlled measurement of LLM-as-judge reliability on a genuinely subjective labelling task: a four-class codebook for how workplace practitioners frame LLM agency, applied to 315 interview utterances from 35 practitioners in India and the UK, with human inter-rater agreement of $\kappa = 0.565$. Four judges (Claude Sonnet 4.5, GPT-4o, Llama-3.3-70B, Qwen2.5-72B) were run through one hosted route with identical prompts and decoding. Judge choice moved macro-$F_1$ by more than 0.36, from 0.751 for the best judge configuration to 0.385 for the worst. The open-weight judges reported mean confidence 0.94--0.95 while scoring accuracy 0.41--0.46, giving expected calibration error of 0.45--0.49 and a high-confidence error rate of 50--58\%: a verifier that is wrong about half the time and says so at 0.9 confidence fails silently. Thresholding on self-reported confidence rescues the frontier judges (Claude: accuracy 0.928 at 51\% coverage) and does nothing for the open-weight ones (accuracy 0.478 while retaining 85\% of predictions). We argue that judge choice is a first-order measurement variable in any pipeline that uses a model to check another model, that judge agreement must be reported against human inter-rater agreement rather than against an assumed ceiling, and that verifier redundancy is worth only as much as verifier independence.
Chat is not available.
Successful Page Load