Beyond Imitation: Evaluating LLM Reviewers by Impact ?
Abstract
Evaluations of LLM reviewers typically ask whether they reproduce human judgments. We instead ask whether LLM review systems can identify papers with greater realized impact, measured using subsequent citations. Across 4,567 ICLR submissions from 2018--2020, an LLM council approach performs comparably to historical human selection on average portfolio impact, improves median impact and high-impact recall, and provides a moderately correlated, nonredundant signal. We examine two central threats to this retrospective comparison---the citation premium from venue acceptance and memorized hindsight---using a regression discontinuity design and a preliminary post-knowledge-cutoff analysis of ICLR 2025 papers. The venue premium favors historical decisions, while the council retains a smaller ranking advantage in the preliminary 2025 sample. These results suggest that LLM councils may be most useful as an independent signal of potential impact, especially at the acceptance margin.