Detecting Research Sabotage from Agent Behavior
Abstract
How can we detect a research agent that deliberately produces misleading work? We study three approaches: judging its work, checking consistency across attempts, and asking follow-up questions. Consistency checks developed for sandbagging also detect instructed sabotage in programming and feature-interpretation tasks. What they compare matters: on APPS, code is more informative than self-explanations, while external summaries retain weaker signal. In machine-learning competition runs, attackers warned about monitoring are easier to catch than unwarned attackers. Generating new actor attempts improves APPS monitoring more than repeatedly judging one attempt. On feature-interpretation tasks, eight follow-up questions do not improve on monitoring the initial answer with its reasoning summary. Two contrasting results sharpen these findings: actor resampling does not improve Ctrl-Alt-Deceit recall at a matched false-positive rate, and a stronger judge improves SleightBench monitoring. More computation is therefore not one intervention: its value depends on whether it produces new behaviour, re-examines the same work, or improves the evaluator.