Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey Study at ICML 2026 Involving Over 24,000 Papers and 17,000 Reviewers
Abstract
Large language models (LLMs) are increasingly used in scientific peer review, yet little is known about how reviewers use them in practice, whether they comply with conference policies, or how such policies affect reviews and paper outcomes. We investigate these questions through a large-scale randomized experiment and post-survey at ICML 2026, in collaboration with the conference organizers. The experiment was embedded in the conference's peer-review process involving 24,661 papers, 17,886 reviewers, and 76,159 authors. All reviewers were assigned to either a conservative policy prohibiting LLM use or a permissive policy allowing limited assistance, with policy assignment randomized for subsets of papers and reviewers. In the randomized comparisons, we found no evidence that policy assignment affects paper scores, reviewer confidence, or final paper decisions; however, reviews produced under the permissive policy were 5.5-7% longer. In the post-survey responses from 1,486 respondents, we found substantial variation in reviewers' experiences and attitudes toward LLM use. While LLM users generally reported improved review quality and reduced time spent, many non-users expressed little need for LLM assistance or raised principled objections. We also observed substantial policy noncompliance: 22.5% of reviewers under the conservative policy reported using an LLM despite a blanket prohibition, and 36.5% of reviewers under the permissive policy reported at least one explicitly disallowed use. Together, this work provides one of the largest empirical examinations of LLM use in peer review and offers guidance for future policy and tool design.