ReviewShield: Defending LLM Reviewers Against In-Paper Prompt Injection under Instruction-Content Entanglement
Somnath Luitel ⋅ Prabhjot Singh ⋅ Suraj Thapa ⋅ Manmeet Singh ⋅ Josh Durkee
Abstract
Large language models are increasingly used to help review the rising volume of conference submissions, but an author who controls their own manuscript can embed hidden instructions designed to manipulate whichever model reads it. This attack is already documented on real preprint servers, and it is harder than generic prompt injection because of *instruction-content entanglement*: the manuscript is simultaneously the untrusted channel carrying the attack and the legitimate object the reviewer must engage with critically, so a defense cannot simply learn to ignore instruction-like text without also degrading the review. We introduce ReviewShield-Bench, built from ~800 open-access arXiv manuscripts and five in-paper injection families, with two held out to test generalization, a 50-paper held-out test set per family, and planted, objectively verifiable flaws that measure whether a defended reviewer still catches real problems. We compare DPO training, a one-line defensive prompt, and supervised fine-tuning (SFT) on preferred responses alone, using three reviewer models in complementary evaluation settings. We evaluate these defenses with a fully paired difference-in-differences design that isolates each defense's effect on attack-induced score inflation from confounds in clean-paper calibration, a confound that makes the standard relative attack-success-rate metric unreliable. Under this corrected metric, DPO at its default learning rate produces no significant reduction in attack effect on any of five families — its strongest defensive-direction result reaches only $p=0.070$ — while one family instead shows a nominal wrong-direction shift ($p=0.039$) that may not survive correction for multiple comparisons. The same null result appears on a second open-weight reviewer model, while a higher learning rate recovers significance on four families, indicating hyperparameter sensitivity rather than a failure of the DPO objective itself. In contrast, SFT-only imitation learning significantly reduces attacker-targeted score inflation relative to default DPO across all five seen and held-out families ($p \leq 4.5\times10^{-3}$), with consistent improvements over the undefended baseline across independent training seeds and without a learning-rate search. However, SFT achieves this partly by *reversing* the direction of the score response rather than eliminating it: on three of five families, the absolute displacement $|\text{AttackEffect}|$ is larger under SFT than under the undefended baseline, meaning injected manuscripts receive an injection-triggered score *penalty* rather than true score invariance. A simple defensive prompt also provides strong protection, particularly against direct override and stealth attacks. We release the benchmark, evaluation harness, and paired difference-in-differences protocol to support more reliable evaluation of future defenses against this emerging threat.
Chat is not available.
Successful Page Load