Who’s Doing the Reviewing? Tracing AI Delegation in Scientific Peer Review
Abstract
Generative artificial intelligence is already being used in scientific peer review. Current work has focused largely on the reviews that AI helps produce, leaving much less understood about the human–AI process that produces them: how researchers consult, evaluate, adopt, and integrate AI assistance while reviewing a paper. This paper traces how 65 researchers with prior peer-review experience evaluated a paper while using an embedded AI assistant. We recorded AI queries, transfers of AI-generated text into their reviews, subsequent editing, engagement with the source paper, and review content. 44 participants used AI, and 30 asked it to generate review content. Reviewers who asked AI to provide review content but did not copy that into their reviews spent about as much focused time with the paper as participants who did not use AI. By contrast, reviewers who copy-pasted AI-generated content into their reviews spent about half as much time viewing the source paper as reviewers who requested generated content but did not copy it. Among these reviewers who copy-pasted AI-generated review content, the pasted text was typically left nearly unchanged, and their reviews were more similar to other reviews of the same paper than those of generation users who did not copy. An exploratory blinded LLM evaluation also found weaker evidence grounding among the reviewers who showed the strongest AI delegation. These patterns suggest that a meaningful subset of reviewers directly adopted AI-generated judgments with little modification, while spending less time with the source paper and producing more similar reviews, and that convergence cuts against a core purpose of peer review: independent experts bringing distinct judgments to the same paper.