Chain of Strategies: Diversity-Preserving Attack Generation for Automated LLM Red-Teaming
Abstract
Diverse attack generation is critical for comprehensive vulnerability discovery in LLMs: safety fine-tuning teaches models to recognize established adversarial patterns, so a red-teaming system that collapses onto a narrow set of attacks quickly becomes obsolete. One paradigm for diversity is a two-stage design that first selects an obfuscation strategy — e.g. via a contextual multi-armed bandit (CMAB) — then generates an attack conditioned on it (Zymet et al., 2026). We introduce Chain of Strategies (CoS), a single-stage framework in which the attacker generates a strategy chain and an adversarial attack string in one autoregressive pass, optimized by the same reward signal. We train CoS with the Trajectory Balance objective of Generative Flow Networks (GFlowNets) (Bengio et al., 2021; Malkin et al., 2022) on top of a chain-of-thought supervised warm-start, using a dense continuous jailbreak reward. On 256 held-out queries against Llama-3.1-8B-Instruct, CoS improves attack success rate (ASR) over the supervised fine-tuned model at every trial budget K ∈ {1,3,5,10} (e.g. 54.7% vs. 33.2% at K=10) while keeping diversity comparable (80.6% vs. 82.1%). Against the two-stage CMAB baseline, CoS already reaches comparable ASR after only a few hundred GFlowNet training steps, at comparable diversity; we attribute the preserved diversity to the reward-proportional sampling behavior of the GFlowNet objective rather than to any explicit diversity reward.