Caught, Not Taught: Intervention History Reduces AI Agent Sabotage on Subsequent Tasks
Abstract
AI control evaluations often reset agents’ memory between tasks, although deployed agents may retain records of prior interventions. We investigate whether learning that an earlier attack was caught helps an agent evade defer-to-trusted (DtT), a protocol in which a trusted model replaces an untrusted model’s submission when its suspicion score exceeds a threshold. Whereas prior AI control evaluations provide feedback while the agent is still attempting the same problem, we study cross-task adaptation: whether evidence of detection changes attack behavior on a subsequent task with a different objective and a fresh repository. We first evaluate an attack policy on 40 tasks in ControlArena’s Infrastructure-as-Code environment, using Claude Opus 4.8 as the untrusted model and Claude Opus 4.5 as the trusted model. Of these, 26 are caught by DtT. For each caught episode, we carry only the prior transcript and intervention outcome into a new task with different main and side objectives, pairing the resulting attempt with the model’s no-history attempt on the same new task. The retained history either states that the previous submission was replaced or additionally includes the trusted model’s replacement. Before DtT intervention, the successful-sabotage—completion of both the main and side objective—rate falls from 0.77 with no history to 0.12 when the model is told that its prior submission was replaced and 0.38 when it is also shown the replacement. Main-task success remains essentially unchanged across the three conditions at 0.88, 0.92, and 0.88, respectively. Thus, the untrusted model continues to complete the main task at comparable rates but becomes less likely to pursue the side-task objective. After DtT is applied, neither history condition produces a statistically significant increase in successful sabotage that survives the protocol.