Koala Science: A Platform for AI Reviewers in the Wild
Abstract
As scientific peer review comes under increasing strain, there is growing interest in using AI to assist reviewers. However, current AI reviewing systems often align weakly with human judgments, produce redundant critiques, and can be gamed through superficial paper edits or prompt injection. These limitations, however, have been studied in settings where AI reviewing systems operate independently and cannot interact with one another. We introduce Koala Science, a platform where independently designed AI reviewing agents discuss the merits and weaknesses of a paper to reach an assessment on the paper's quality. We evaluated Koala Science in a seven-day competition in which 46 AI agents reviewed 347 papers concurrently under review at ICML 2026. Koala Science scores were more aligned with ICML outcomes than simple LLM-reviewer baselines, while remaining below a heavily engineered AI review system. We annotate a subset of the arguments present in these reviews to assess their quality, finding that 40% of them are categorized as relevant and verified by both annotators. Extrapolating this rate across the platform yields an estimated 30.2 relevant and verifiable arguments per paper, comparable in count to the 29.5 arguments extracted from human ICML reviews on the same papers. Koala Science also produced more distinct arguments than what previous critiques of AI-assisted review had shown. From these results, we develop four pillars that AI-assisted peer review should attempt to follow, which guide our future iterations of the platform.