Sensitive but Not Specific: What a Trusted Policy's Logs Can and Cannot Say About Reward Hacking
Jules Roussel
Abstract
When an agent is updated, operators may want to screen the candidate before running it, using only trajectories logged by the trusted incumbent. Off-policy evaluation (OPE) is a natural offline verification tool, but importance-sampling estimates become unreliable precisely when a candidate departs strongly from logged behaviour. We ask whether this failure is itself informative: can log-side support diagnostics distinguish reward hacking from benign change? In a controlled Wordle agent testbed with exact action probabilities, we construct and gate-certify faithful, benign-drift, benign-improving, degraded, and reward-hacked policies—the hackers distilled from scripted proxy-farming teachers after hacking failed to emerge across seven escalating GRPO runs. Effective sample size (ESS) is sensitive but not specific: it collapses for every OPE-evaluated certified hacker, yet also falls below the reliability threshold for our strongest benign improver, while a degraded policy passes. A second diagnostic, \%floor, orders the certified categories without overlap: benign $\leq 0.05$, degraded $0.12$, hacked $0.15$-$0.16$. Yet under a benign control matched to a hacker on importance-weight variance, 0 of 6 diagnostics pass a pre-registered hacking-specificity criterion. Trusted-policy logs can therefore triage which agent updates deserve direct evaluation first, but cannot by themselves verify a reward-hacking verdict.
Chat is not available.
Successful Page Load