A Same-Generator Corpus and Classifier Floors for AI-Control Monitor Evaluation
Abstract
Trusted monitors for AI control are tuned and ranked on stored pairs of code: for each programming problem, an honest solution beside a solution carrying a planted backdoor. In the corpora used for this, the honest solution is human-written and the backdoor is written by a language model, so the two classes differ in who wrote them as well as in whether they are sabotaged. Filtering to working solutions and stripping every comment does not remove that difference. On the filtered, comment-stripped split, a bag-of-words classifier separates the two classes at AUROC 0.806, and a classifier that reads no code at all, only problem metadata, reaches 0.772. The signal transfers across backdoor generators and is almost as strong on backdoors that never work, so it is not a property of the sabotage. We then rebuild the corpus so that one model writes both classes on the same problems, and release it. Holding the writer constant removes between a third and nine-tenths of the separability, and holding the prompt constant as well removes most of the rest. An LLM monitor moves the same way, from 0.958 on stored pairs to 0.566 on same-generator pairs. The label installs a second artifact: a problem whose backdoor fails contributes its human honest solution, so the negative class records attack failure. Stored-pair evaluations should report a no-code and a bag-of-words floor against permutation nulls, generate both classes from one model under one prompt, and define the negative class by honesty rather than by attack failure.