Attacking the Oversight, Not the Mission: Backdoor Evasion Against Activation-Based Monitors in Reasoning Agents
Abstract
Reasoning models that think in a continuous latent trajectory rather than emitted text remove the one property chain-of-thought monitoring depends on: a token stream a monitor can read. We ask what happens when a monitor built on the resulting activations is attacked directly. We plant a backdoor in a recurrent continuous reasoner, confirm a linear probe detects its internal state on the latent trajectory, and adversarially fine-tune the model to push that probe toward chance, matched step-for-step against a control that receives no such pressure and holds behavior and clean utility fixed. Across a full λgrid and three seeds, the attack produces real but narrowly localized evasion: separability collapses only at the exact latent position the loss term targets, is fully recovered by a fresh monitor probing anywhere else in the (redundant, untouched) trajectory, and never rises elsewhere – destruction without relocation. The identical attack against a capacity-matched ordinary (explicit chain-of-thought) model produces zero detectable evasion at any strength, even against its own literal optimization target. We traced this to a baseline mismatch between the features the attack can optimize and the features a monitor actually observes, then tested the obvious fix directly by training the attack on the right features – it changes nothing, so the mismatch is not the reason this arm resists evasion. Repeating the entire protocol at two further, independently trained scales reproduces both results. Two genuine exceptions are worth reporting rather than smoothing over: the localized drop’s extent widens at the intermediate scale, and whether the frozen-target attack succeeds at all becomes seed-dependent at the largest one, a heterogeneity we confirmed was not an artifact of an underpowered probe before treating it as real. A claim this specific has to be able to break; here, in two checkable ways, it did, in a way that bounds how far a single-scale probing result should be trusted to generalize.