When the Judge Intervenes: Instruction Injection and Evidence Poisoning in Sequential LLM Pipelines
Abstract
We study whether error correction in sequential multi-agent pipelines depends on the kind of adversarial input. In our setup, we attack a Planner →Worker → Judge pipeline through three input channels: an injected instruction (GSM8K), a false conclusion added beside intact evidence (HotpotQA), and a corrupted fact that the answer depends on (HotpotQA). Across two model families (Claude Sonnet 4.6 and GPT-5.4), we examine each agent's output separately to measure whether the judge removes, preserves, or introduces the attacker's answer. We find an asymmetry across channels: under instruction injection the worker adopted the attacker's answer in nearly every run and the judge removed it in 64 of 90 such runs (Sonnet) and 87 of 89 (GPT), while under evidence poisoning the worker adopted it in 19 of 264 runs and the judge removed none of them (95\% CI upper bound 17\%). In the matching clean conditions the judge changed no worker answer, so revision occurred only under attack. We further find that verification can introduce the very error it exists to prevent: all six revisions the GPT judge made under evidence poisoning replaced a correct worker answer with the poisoned one. Overall, our findings suggest that a judge that removes an injected answer should not be assumed to remove a poisoned one, with verification in one evaluated condition introducing the error it was meant to catch.