Wrong the Same Way: Cross-Prompt Sharing Both Suppresses and Amplifies Reward Hacking in RLVR
Abstract
Two accounts of how reinforcement learning with verifiable rewards (RLVR) interacts with an imperfect verifier coexist in the literature: in one, the rate of reward hacking is inherited from the base model rather than learned, while in the other, RL actively amplifies it. We show that these are one mechanism on opposite sides of a threshold, separated by a property of the verifier's errors that neither account varies: the universality of a leak, or the fraction of prompts that accept the same wrong response. At its KL-regularized optimum, RLVR acts as a filter which reweights the verifier-accepted set as a whole without altering its internal composition, so the hack rate is inherited exactly. Taking that as a null model, we write the exact drift away from it as a sum of covariances between reward and the empirical tangent kernel, which isolates a cross-prompt term that vanishes identically without parameter sharing. Below the threshold, sharing drives a leak below the base rate, which is inheritance and more; above it, the same sharing manufactures amplification. We confirm both signs in a controlled RLVR environment small enough to compute the entire response distribution exactly, in which universality and the static false-positive rate are independent knobs. Replacing the shared architecture with independent per-prompt copies removes the dependency on leak breadth, neutralizing the drift slope. Targeted controls further show that sharing at RL time alone is insufficient: exploitation is amplified only when the base model's representation already encodes the verifier's acceptance pattern. Verifier hardening must therefore be budgeted against the model's pre-existing shared error modes, since patches targeting universal leaks yield significantly higher true accuracy than uniform error reduction.