Top-$k$ Identification with Correlated Biased LLM Judges via Anchor Leverage
Sixiong Xie ⋅ Zhuofan Shi ⋅ Haiyang Shen ⋅ Yun Ma ⋅ Xiang Jing
Abstract
Large-scale model evaluation is increasingly delegated to LLM judges. This makes evaluation cheaper, but it creates a new statistical problem: if several judges share a systematic bias, averaging more judge scores can make a ranking look more confident without making it more correct. We study fixed-confidence top-$k$ identification when judge scores are cheap, correlated, and biased, while trusted anchor labels are costly. Under a low-dimensional factor-bias model, we show that top-$k$ identifiability is controlled not by the total number of anchors, but by whether anchors span the feature directions separating arms across the top-$k$ boundary. This yields an instance-dependent cost characterization with two components: effective information from correlated judge scores and a profiled anchor-leverage term quantifying the cost of removing shared bias. We propose Profile Track-and-Stop, a sequential allocation algorithm that tracks the resulting cost-optimal allocation with a conservative pilot anchoring phase and asymptotically matches the lower-bound constant. Experiments on a 10-judge Arena-Hard-v2.0 evaluation instance with 23,545 real API judge scores, together with synthetic and semi-synthetic ablations, are consistent with the predicted failure mode: anchor-free methods can plateau under cross-family judge bias, whereas boundary-aware anchoring improves top-$k$ recovery at lower cost.
Chat is not available.
Successful Page Load