No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions
Abstract
As AI-generated reviews move from experimental tools into peer-review infrastructure, most robustness concerns have focused on explicit attacks such as hidden instructions and prompt injection. We study a harder and more policy-relevant failure mode: no hidden text, no prompt injection, and no changes to experiments, figures, or numerical results. The attacker modifies only presentation-level content across the full paper, such as the abstract, contribution framing, and discussion. We introduce adversarial repackaging: a closed-loop attack that uses AI-reviewer feedback to search for presentation-level revisions with scientific evidence fixed. Across mainstream AI reviewers, it achieves a 75.1% attack success rate and a mean score gain of +1.21/10. The effect is not explained by ordinary prose polishing: strategies that change how the reviewer interprets the paper, such as related-work repositioning, substantially outperform surface polishing and formatting. Our analysis reveals two deeper structural failure modes. First, AI reviewers are easier to impress than to convince: highlighting strengths reliably increases perceived merit, while attempts to dissolve weaknesses frequently backfire. Second, AI reviewers confuse the appearance of addressing a limitation with actually resolving it, allowing unchanged evidence to be reinterpreted as a stronger scientific contribution. The vulnerability holds across models and review templates, indicating a structural, not single-model, deficiency. These results show that the deployment risk is not only malicious hidden instructions, but the emergence of paper presentation itself as an optimization surface. We release a contamination-free rolling benchmark and attack framework for testing resistance to presentation-only review gaming.