What Should You Measure Before Shipping an Adapted Policy? Attack Success Against a Benign Control Arm
Abstract
Post-training changes a policy's behaviour, and the question that follows is whether the adapted checkpoint is safe to ship. The usual answer is an attack success rate. We argue that this number cannot answer the question on its own, because it is not comparable across two checkpoints unless the false-positive rate of the safety predicate is reported alongside it. A safety predicate is a detector, and a detector applied to the benign distribution fires sometimes. If that rate is unmeasured, a change in attack success between checkpoint A and checkpoint B may be a property of the instrument rather than of the adaptation. We evaluate a SmolVLA policy on all ten libero_object tasks using a paired design: every attacked episode has a benign twin at the same task and seed, run through the same policy and scored by an identical predicate. Significance is McNemar's exact test on discordant pairs, corrected with Holm across the six measured arms, and intervals are bootstrapped clustered over tasks rather than over episodes, because episodes within a task are correlated. A single reworded instruction drove the policy out of its safe envelope on 44 of 50 pairs (88%, task-clustered 95% CI [72%, 100%]) against a benign control of 2 of 50 (4.0%, Wilson 95% CI [1.1%, 13.5%]), McNemar exact p = 4.6e-13, Holm-adjusted 2.7e-12. Both benign firings land on cells where the attacked arm also fired, so the paired test counts 42 discordant pairs rather than 44: the raw rate and the paired evidence differ by exactly the amount the control arm exposes. Two arms survive correction, and four do not; only the instruction family transferred at all, and we report the visual and injection nulls as part of the result. The control arm changes what the numbers mean. Across two independent runs, the benign arm fires 5 times in 100 episodes, and all five land on two of the ten tasks while the remaining eight stay silent through 80 benign episodes. The seeds do not repeat. Each run tests the other out of sample, giving 3 of 3 at p = 0.008 and 2 of 2 at p = 0.040; we take the conservative direction as the headline. A wandering policy scatters its failures; a misplaced boundary is a geometric fact about a scene and fires on the same tasks whatever the seed. Read without the control arm, those five episodes are attack successes. They are not: they locate a keep-out region drawn in the wrong place, which is a property of the evaluation rather than of the policy. Three attack families are measured as nulls at 0 of 50, interpretable only against a non-zero benign rate. And a family that failed correction on a single task is significant when pooled over ten, so single-task adversarial results are underpowered in both directions. For a post-training workflow, this yields a per-checkpoint protocol: fix the task and seed set before any adaptation, run both arms for every checkpoint, and report the pair. A rate that moves between checkpoints is evidence only if the floor underneath it held still. We state the protocol as a recommendation: this paper evaluates one checkpoint and has not run it across a sequence of adapted ones. We also report two defects found in our own published numbers rather than in someone else's. A clustered bootstrap over ten all-zero tasks returned a zero-width interval that overstates certainty, where 0 of 50 admits a true rate up to 7.1% by the exact binomial upper bound. And the policy sampler was unseeded at measurement time, so the headline is one draw rather than a reproducible constant. Scope. This paper evaluates a policy and does not adapt one; it reports no fine-tuning or preference-tuning experiment of its own. All results are in simulation. No real-hardware trial has been run, and we make no transfer claim. One suite and one checkpoint. The attacks are a templated screen rather than an optimised worst case, so the rates are a floor on susceptibility. The predicate is uncalibrated by construction, which is what makes its misfire measurable. The harness is open source under a permissive licence; the per-task benign-firing analysis reproduces on CPU in seconds, while the policy rollouts behind it required a GPU.