Learning Automated Checklists to Reduce Lottery Effects in Scientific Peer Review
Abstract
Reviewer lottery error describes error due to a small set of peer reviews attending to only some of the issues with a paper that a larger pool of eligible reviewers would raise. We show this error can be reduced by using a historical review corpora to learn a coarsened checklist of criticisms that can be checked against the paper, which can be surfaced for reviewers. Our approach selects the granularity of the coarsening to maximize the disparity between the share of overlapping signals raised by two reviewers of the same paper versus of different papers, guarding against an overly generic checklist or an overly specific one. We validate the method through a LLM-simulated peer review study, where the review target is the score assigned by a large number of reviewers but review assignments are much smaller. Using NeurIPS 2021 data, we find that our coarsening selection approach identifies the checklist that best balances reduction in reviewer lottery error against added bias from the shared checklist representation: among the five granularities we tested, it has the lowest error, and it reduces error relative to unaided review for committees of fewer than six reviewers.