Stress-Testing the Truth Advantage in Language Model Debates
Abstract
As AI systems become more capable, evaluation becomes difficult when the evaluator cannot independently verify the model's answer. This is the scalable oversight problem. AI debate addresses it by placing two language models in opposition before a weaker judge, testing whether the truthful side has a systematic advantage over the deceptive side, a property referred to as the Truth Advantage. In this setting the judge is the verifier, and the strength of the protocol rests entirely on how reliable that verifier is. Prior empirical work supports the Truth Advantage under standard benchmark conditions, leaving open whether it holds when question wording or structure changes while the correct answer remains fixed. We establish consultancy and debate baselines on the QuALITY reading-comprehension benchmark using two independent judge models. Consultancy judge accuracy averages 57% across the two judge models; debate accuracy averages 80%, a difference reflecting the structural contribution of the opposing debater. We then stress-test the Truth Advantage through a three-stage Quality-Diversity adversarial search targeting the subset of questions both judges answered correctly at debate baseline, a pool where debate-baseline accuracy is 100% by construction. Across three stages, 73.2% of this pool (Wilson 95% CI [68.3%, 77.7%]) yields at least one framing under which both judge models return an incorrect verdict on a full debate transcript. Projected back onto the full protocol comparison, these pool-level failures give a worst-case adversarial judge accuracy of approximately 30%. These findings indicate that question framing constitutes a reproducible vulnerability in debate-based oversight protocols and motivate adversarial framing stress-tests as an important complement to standard debate evaluation.