Reward Hacking as a Transmissible Strategy in Multi-Agent Coding Systems
Abstract
Multi-agent systems are increasingly used to solve coding tasks by coordinating agents that exchange code, memory, and plans. Coding agents used in these systems are often optimized with reinforcement learning (RL) based on feedback from automated evaluators. Yet passing an evaluator need not require a correct solution, creating opportunities for reward hacking. Prior work has largely treated reward hacking as a single-agent failure. In this work, we show that reward hacking can become a system-level failure in multi-agent coding systems: a strategy elicited in one agent can propagate to other agents and alter system coordination. Across four open-source backbones, we study two knowledge settings. The prompted setting explicitly describes the mechanisms, whereas the mid-training setting models exposure to facts about the coding environment during upstream training. Coding RL elicits and amplifies reward hacking in both settings. Using post-RL models from the mid-training setting as source agents, we find that their strategies propagate through code, shared memory, and natural-language plans. Even source agents based on small open-source models induce reward hacking in stronger closed-source targets. Moreover, propagation can intensify along a transmission chain after the original source leaves as the code acquires more persuasive justifications. Reward hacking can also bias orchestrator routing toward reward-hacking agents and survive candidate aggregation. Self-review and verifier gating reduce but do not eliminate delivered hacks. Thus, reward hacking can spread beyond its source, persist, and reshape multi-agent coordination.