Surface Form or Shared Belief? A Behavioral Test of Whether LLM Judge Agreement Is Diversifiable
Shayan Amir
Abstract
Individual judges on LLM-as-a-judge panels provide correlated signals about the same underlying fact, which results in diminishing information gains as panel size increases. We ask whether agreement among LLM judges arises from shared wording or from a shared evaluative belief. We hold a 6-judge panel of frontier LLMs from 2 model families and vary only the view of each item, comparing meaning-preserving rewrites with identical input and shared paraphrasing on ChaosNLI-MNLI and RewardBench Chat/Chat Hard. If wording drives agreement, distinct views should raise the panel's effective sample size by at least $+0.5$ votes and outperform a noise-matched control. They do not. The gains are $+0.12$ and $-0.02$ on the tasks, respectively, both consistent with zero. Of the 47 items that all six judges answered incorrectly at baseline, 26 remain unanimous errors after rewording, and all 7 unanimous preference failures persist. Our work indicates that, for this panel and these tasks, agreement is not diversifiable by changing surface form and is better read as a convergent belief about the item's meaning.
Chat is not available.
Successful Page Load