Does Adversarial Diversity Survive a Guardrail? A RainbowPlus Extension
Abstract
Quality-diversity (QD) red-teaming methods such as RainbowPlus are evaluated almost exclusively against undefended target models, yet real deployments typically layer a cheap system-prompt safety instruction on top of the base model. We extend RainbowPlus with a controlled comparison of a bare Qwen2.5-7B-Instruct target against the identical model wrapped with a generic system-prompt guardrail, repeating each condition across six random seeds. The guardrail lowers mean attack success (79.6% to 48.5% ASR) and mean archive coverage (62.8% to 15.8% of cells filled), but its most striking effect is on stability: guarded-condition ASR ranges from 16.7% to 90.9% across seeds (standard deviation 24.5, roughly 13x the bare condition's), meaning a single evaluation run under a guardrail is a poor guide to how well it actually works. This argues for both diversity-aware and repeated-run evaluation when measuring robustness interventions, rather than trusting a single scalar from a single run.