When can we trust untrusted monitoring? An AI control safety case sketch across collusion strategies
Abstract
Potentially misaligned, untrusted, AIs are being deployed as AI agents despite being capable of harm. One in-use approach to reducing risk is untrusted monitoring: using one untrusted AI to oversee another. AI control tries to create deployments that are robust to intentional subversion. The key vulnerability of untrusted monitoring is collusion: the monitor could behave honestly on test inputs, but wave through malicious outputs it recognises as coming from another instance of itself. In this paper, we ask: could we show that a given untrusted monitoring deployment is robust? We sketch a safety case that clearly presents the strongest argument for safety -- improving the state of the art empirical methodology whilst highlighting necessary assumptions and significant unsolved challenges. The core of our safety case is a bound on the untrusted monitor's ability to collude by recognising its own outputs. We develop a taxonomy of collusion strategies -- passive self-recognition, causal, acausal, and combined -- relaxing the assumption in prior work that colluding AI instances can share a secret signal in advance. We introduce an SFT-based elicitation method to conservatively upper-bound passive self-recognition. On APPS and BigCodeBench with GPT-4.1, we find that elicited passive self-recognition outperforms the best previously studied collusion strategy, achieving a red team win rate in APPS of 9.8\% versus 2.8\% at full usefulness under summarisation, the most effective signal jamming defence.