CouncilAttack: Red Teaming LLMs with a Deliberative Council over Fabricated Multi-Turn Dialogues
Abstract
Automated red teaming is widely used to surface LLM jailbreaks before deployment, yet most automated attackers draw their proposals from a single model. We present CouncilAttack, a closed-loop red-teaming framework in which a council of peer attacker LLMs independently generate candidate fabricated multi-turn dialogues, cross-critique and rank them, aggregate first-place preferences using Borda tie-breaking, and pass the ranked output to a chairman that refines or combines it into the attack sent to the target, with an in-loop 0–5 harm judge steering the next turn. To the best of our knowledge, this is the first attacker-side red-teaming framework to aggregate preferences across peer proposers through an explicit social-choice rule. Evaluated on HarmBench against four open-weight targets under a single fixed judge, a single configuration reaches up to 96% ASR≥4 and a five-configuration portfolio reaches 98–100% across all four targets, outperforming strong single- and multi-turn baselines under the same judge; the maximum-harm attacks then transfer verbatim to seven closed frontier models, reaching 58.8% ASR≥4 on gpt-5. At a matched target-query budget, a same-model council significantly improves strict full-jailbreak success over that model attacking alone (45.9 vs 37.0% ASR≥5, p=0.018).