Do LLM Reviewers Correct for Multiple Comparisons?
Abstract
Conferences are beginning to deploy language models as reviewing aids, but it is unclear whether these models propagate statistical concerns into their conclusions. We construct a controlled benchmark in which the reported result is held fixed at 285/500 correct predictions (57.0%, exact two-sided binomial p = 0.002) and only the disclosed number of evaluated configurations N varies, from N=1 to N=10,000. Across 157 reviews from two open-weight model families, spontaneous detection is close to binary. Flagging averages 66% across every N >= 5 with no monotone trend; a template-clustered regression on log10(N) gives a slope of +0.14 (95% CI [-0.24, +0.52], p = 0.46), so we exclude a strong response to search size but not a moderate one. Measured against the reference itself, which steps from "evidence stands" to "evidence fails" at N=25, reviewers are misaligned in both directions: they criticise the evidence in 58% of reviews below the crossover, where it survives correction, and let it stand in 29% above. The observed step across the crossover is +0.12 (95% CI [-0.01, +0.36]) against a reference step of +1.00. No review named a correction procedure or computed an adjusted threshold. Yet asked directly whether the evidence survives the disclosed search, the same models answer correctly in 42 of 42 cases above the reference crossover. The failure is therefore one of application rather than ability: these models reach the correct multiplicity-aware conclusion on request and never do so unprompted.