When Vulnerability Is Not Exploitation: Reward Hacking in Coding Agents
Abstract
Reward hacking is often treated as a single failure mode in which an agent exploits imperfections in an evaluation mechanism rather than completing the intended task. In agentic coding environments, however, environment vulnerability and agent behavior are distinct: the same structural weakness may coexist with honest task completion, effort-substituting shortcuts, or deliberate exploitation. We study this distinction at both the environment and trajectory levels. We analyze 3,009 environment-review findings to develop an intervention-based taxonomy of structural vulnerabilities, and 5,343 trajectories from six model variants across anonymized machine-learning engineering environments to distinguish honest behavior, shortcut behavior, and executed reward hacking. We find that structural vulnerability does not directly translate into exploitation. Shortcut behavior occurs in 17.05\% of trajectories, compared with 0.43\% for executed reward hacking, and structurally exploitable vulnerability categories do not necessarily exhibit high exploitation rates. Model differences are also substantially stronger for shortcuts, while executed reward hacking remains rare and less stable across annotators. These results suggest that environment vulnerability, effort substitution, and deliberate reward exploitation should be measured separately rather than collapsed into a single ``reward-hacking rate.''