A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
Abstract
AI reviewers increasingly support consequential scientific evaluation, yet content-preserving changes in wording can alter their judgments. As LLMs make manuscript rewriting inexpensive, this sensitivity can reward rhetorical optimization over scientific improvement. We argue that Rhetorical Robustness is an important but neglected requirement for trustworthy AI reviewers, jointly requiring within-paper stability across content-preserving rewrites and between-paper discrimination of scientific differences, with human alignment evaluated separately. We introduce RobustReview, a controlled full-manuscript benchmark comprising 60 anonymized ICLR 2026 submissions, 10 rhetorical conditions, two rewrite producers, and 1,260 manuscripts. Seven metrics evaluate 30 configurations spanning general-purpose, specialized, and agentic review systems. The benchmark reveals false robustness, in which low rewrite sensitivity coincides with score collapse across papers. No configuration is consistently robust, and human alignment ranks reviewers differently from rhetorical robustness. We further introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a content-normalized judgment obtained by extracting and reviewing a structured science core. By complementing manuscript review rather than replacing it, this design seeks greater rhetorical stability while retaining manuscript-level assessment. SciCore achieves the strongest joint stability-discrimination profile in our primary comparison while maintaining competitive human alignment. Together, RobustReview and SciCore establish rhetorical robustness as a distinct evaluation target and show that structured scientific content can support more trustworthy AI review.