LLMs Learn to Evade Latent Monitors from Feedback Alone
Abstract
Latent space monitors promise to detect undesired behaviors in LLM agents by inspecting their internal activations rather than their outputs. However, interactive monitoring creates a feedback channel where each verdict the monitor makes leaks information to the model about how its internal states are being evaluated. We ask whether an agent can exploit this feedback to infer what a monitor detects and selectively edit its activations to evade detection, without other information about the monitor. Surprisingly, off-the-shelf models already produce monitor-specific activation edits aligned with the direction needed for evasion, but at insufficient magnitude. Simply rescaling these edits by a factor of 8 reduces the monitor's TPR @ 1\% FPR to 0.27. A rank-1 LoRA amplifies this behavior into effective evasion within the forward pass, reducing TPR further to 0.04 on held-out concept monitors while leaving other concepts at their normal detection rates. Capabilities on standard benchmarks are retained under this finetuning, and the evasion skill survives monitors retrained on the new activations. Mechanistically, we find evidence that the model computes its activation edit from the prior in-context turns, and show that its edit becomes more aligned with the monitored direction as more examples are provided. These results demonstrate feedback-conditioned control over activations and suggest that latent monitoring should be treated as an interactive process where agents can observe and respond to oversight measures.