A Verifier Validated on a Proxy Can Reverse in Deployment: The Signed Effect of Audit Certainty
Abstract
We calibrated a verifier on a proxy task, deployed it unchanged, and its strongest setting became its worst: raising audit certainty — the lever that cut violations on the proxy — raised them on deployment (0.433→0.800; a no-penalty control shows the verifier still helps, so what inverts is its certainty dial, not its net value). The claim in one sentence: proxy calibration of audit certainty need not transfer across temporal structures for learning agents. Two findings support it, with different standing. The first is derived — exactly for the variance, under a threshold approximation for the behavioural sign. Classical deterrence says a verifier's budget matters only through the expected penalty E=pc (audit frequency × severity); for an agent that learns from the signal it does not, because rare severe penalties inflate the variance of the learned violation value. We prove an exact stationary-variance result and, under a threshold approximation, derive from it a signed certainty law — sign(∂P/∂p)=sign(Δ), with Δ the violation's value advantage — whose directional prediction, made before the experiment, is that the lever must invert rather than merely weaken; the experiments confirm it. The second is causally isolated but not derived: the threshold's location moves with task timing. Equalised, it sits at -0.007, on the variance-channel prediction; delaying honest reward alone pushes it past every value gap we can measure (beyond -0.45; timing contrast z=12.1), so an entire band of value gaps changes sign and a proxy whose timing differs from deployment can certify exactly the wrong policy. Hence the design rule: calibrate on matched timing, or measure Δ_γ under deployment timing before trusting a proxy's verdict — and below the boundary spend the marginal oversight dollar on detection, because above it the same dollar buys more violations.