Benchmarking Agentic Review Systems
Abstract
A new class of agentic review systems is emerging as a remedy to the pressure placed on peer review by AI-assisted research, but it is unclear how they should be evaluated. We evaluate two open-source systems (OpenAIReview and coarse), one proprietary system (Reviewer3), and a zero-shot baseline, across six LLMs spanning frontier and efficient models. First, we test whether AI reviews separate weak from strong ICLR/NeurIPS papers, as approximated by citations and acceptance decisions. Every system performs above chance in pairwise accuracy, and the best is OpenAIReview + GPT-5.5 at 83.0%. On the same papers, we rate comment precision with six LLM judges and human validation, and the best system attains 89.3%. Second, to test whether systems catch errors with known ground truth, we inject four categories of errors into papers across eight arXiv subject classes and measure detection recall. The strongest configuration (OpenAIReview + GPT-5.5) catches 71.6% of injected errors, leaving substantial room for improvement. Model-controlled comparisons show that recall depends on both model strength and harness design, while the union across six models reaches 83.3%. Together, by evaluating full review systems backed by state-of-the-art models on real papers, we show that AI reviews, while imperfect, can already separate weak from strong papers, flag genuine issues, and catch important errors.