Too Early for AI-Assisted Peer Review: A Systematic Account of the Limits and Opportunities of Automating Human Judgment
Abstract
OpenReview's October 2025 announcement to pilot an AI-assisted peer review process for AAAI-26, and the news in November 2025 that ICLR-26 was flooded with AI-generated reviews, show that a discussion about how to design the peer-review process of the future is long overdue. To foster a community-wide debate, we distill a set of 24 core capabilities required for scientific review as the foundation for a taxonomy on the dimensions and tasks relevant to peer reviewing. On this basis, we systematically and critically reflect on whether AI assistance tools should be integrated in the review process. We argue that the current state of the art does not warrant a general inclusion of such technology: Apart from security issues and ethical concerns, the capabilities of generative AI are for the most part not reliable enough yet, insufficiently tested, or inadequate by design to be integrated in the peer review process. With this paper we pave the way toward structured benchmarking across multiple dimensions.