The Grade Says Little, the Score Orders: AI Reviewer Panels Against a Venue's Ratings and Decisions
Jiecheng Lu ⋅ Shihao Yang
Abstract
Authors already run language models over their drafts, while venues are beginning to use them as reviewers. We ask what that output measures against a venue that knows which it accepted. Across 2,972 reviews of ICLR 2026 submissions, a frozen five-model panel recovers the venue's post-rebuttal ordering of rejected papers at $\rho=+0.595$ (AUC 0.820). Its 480 baseline readings, spanning the decision, award the top grade once. Its bar for a recommendation to accept is Minor Revision, the prompt's borderline accept at an ICLR rating of 6, which just 56 of those 480 reviews reach. Scoring every paper twice shows re-run noise exceeds the median 0.115$z$ gap between adjacent reference-field papers, so even the ranking is coarse. An eight-model, six-vendor roster preserves the ordering but relocates the grade, whose absolute level is the panel's, not the paper's. Reviewer identity explains 8.4\% of score variance in the wide-range population but 18.8\% in the homogeneous one a committee faces. Averaging against the mean member shows a gain too small to measure: +0.043, reaching +0.088. Joining sentences moves the composite 0.86 of its noise band; relocating figures leaves it at 0.13, reaching its licensed channels. Every arm moves the grade more than a re-run. Agent review output is therefore usable as a calibrated relative ranking across a wide quality gap, unusable as an absolute grade, and only after a presentation-invariance audit. We release the instrument, reviews, calibration parameters, transforms, and the 19,814-submission ICLR frame.
Chat is not available.
Successful Page Load