The Objective, Not the Corpus: Where the Training Signal Goes in Automated Peer Review
Abstract
Automated reviewers are increasingly built by fine-tuning a language model on a corpus of real peer reviews. We ask what that training actually buys. On 781 held-out ICLR papers with known outcomes, and giving every method only a paper's title and abstract, fine-tuning Qwen2.5-3B on 28,254 human reviews raises accept/reject discrimination from AUROC 0.473 to 0.582, a real gain (p < 0.001, DeLong). But the result is statistically indistinguishable from simply prompting an off-the-shelf instruction-tuned model of the same size (0.566, p = 0.13 to 0.85), and it is beaten by a bag-of-words classifier tuned over 96 configurations on the same input (0.696, p <= 0.001). Holding the model, the data and the input fixed and replacing next-token generation with a scalar regression head gives 0.736: a large gain over generative training (p < 0.0001) that nonetheless only reaches parity with that linear baseline (p = 0.013 to 0.066). Where the training signal falls therefore matters more than which corpus it comes from, and still does not lift the model past what a linear model extracts from the same input. The mechanism is dilution: the rating occupies one of roughly 838 tokens in a training example. We then train such reviewers repeatedly on their own output. Across seven contamination levels, three seeds and six rounds, the vocabulary of generated reviews is unchanged at 1% contamination, falls by 37% at 5%, and saturates near -79% by 25%. Contaminated models write longer, not shorter, reviews, so we also measure with every review truncated to a common length: the decline persists but its size depends on the window, at -41%, -25% and -8% for budgets of 200, 120 and 60 tokens. Two defences requiring no ability to detect machine text rank differently depending on which property of a review is measured: refreshing half the corpus with human reviews each round preserves 76% of a review's section structure but only 6% of its vocabulary variety, while retaining all past data recovers 117% and 28% of those same two.