When Verification Becomes a Training Signal: Reward-Hacking Propagation in Multi-Agent Systems
Rishikesh K Muralimohan ⋅ Yash A Patil ⋅ Vatsal Desai ⋅ Cuong H Nguyen ⋅ Diksha Shrivastava
Abstract
When language-model agents work together, they often see one another's code and an evaluator's PASS/FAIL verdicts. That shared context helps coordination, but it can also give a successful shortcut a way to spread. We ask whether one agent that has been conditioned to game an evaluator can change the behavior of teammates that were never told to do so. Our experiments use six-agent coding networks and ImpossibleBench tasks whose written requirements conflict with their tests. We call the specially conditioned seed agent \emph{Patient Zero}. It privately receives verified reward-hacking examples and a trigger instruction; the other five agents never see that private material. We compare each Patient-Zero network with a matched network whose seed is conditioned to behave honestly, while keeping the tasks, models, communication policy, verifier, and peer prompts the same. Across 90 matched pairs, Patient Zero raises post-exposure adoption by $13.8$% (95% CI $[10.4,17.1]$) and successful hacking by $9.1$%$[6.7,11.8]$. The effect appears under every communication policy, although successful execution depends strongly on the receiver model. In a separate 728-call follow-up, the behavior carries over to unseen impossible tasks when an adopter keeps its own earlier history, but we do not see corresponding broad clean-task drift. The takeaway is straightforward: verifier outputs do not only score an agent. Once they enter shared context, they can also influence what other agents do. Reproducibility: Note that our code will be made public upon acceptance.
Chat is not available.
Successful Page Load