What the Monitor Saw: activations, pooling, and evidence combination in AI control cascades
Abstract
Deployed oversight of AI agents is often a cascade; a cheap monitor watches each action, and escalates those it finds suspicious to a stronger monitor. We study cascade design choices on LinuxArena, a public agentic sabotage benchmark, and find that the way we use evidence dominates. If we replace the cheap monitor's verdict with the strong monitor's - the usual way to build a cascade - a better readout of the cheap monitor's activations buys nothing. If we combine the two verdicts instead, those same activations roughly double the routed catch. Under this combining rule, our main result is free to deploy - a 64-unit MLP head reading the cheap monitor's residual stream, making no strong-monitor calls at all, matches what paying for the strong monitor on every call buys, scored through the benchmark's own evaluation code and under a stricter trajectory-level false-positive convention of our own. We find that what that head reads is a signal the cheap monitor has, but does not print - a linear probe on the same activations recovers 1.6x as many of the monitor's wrong calls as its own verbalised confusion does, and an ablation rules out the printed suspicion score along with the other uncertainty signals the monitor can supply about itself. Which signal routes best is not intrinsic to it either; it flips with how per-call evidence pools into a trajectory verdict, and the flip recurs in a separately fitted estimator, where a context gain of +0.43 catch under MAX collapses to +0.02 under SUM. Finally, what the strong monitor is shown matters as much as when it is called; disclosing the cheap monitor's verdict, the natural deployment default, actually costs safety.