Misalignment Emerges During Benign Iterative Self-Improvement
Abstract
AI systems can improve themselves, bringing both the promise and danger of leading them towards superhuman capabilities. Self-distillation has gained popularity as one of the main algorithmic techniques to realize such potential by bootstrapping the model's own generations under additional feedback information. In this work, we identify a safety failure of such iterative improvement loops. A safe base model can become misaligned while being trained via self-distillation, even if only trained on its own (initially) safe refusals and with a fixed and non-harmful feedback. Models become compliant across different unseen and unsafe tasks as measured on StrongREJECT and SORRY-Bench, e.g., by instructing the user on how to build an artisanal bomb or cause physical harm. This failure mode persists across sizes and families of open frontier models like Qwen3 and Phi-4, and across both on- and off-policy self-distillation. Additionally, we show that self-distilling starting from safe refusal strictly on harmful prompts is not necessary: a model self-distilling on a mix of safe and unsafe prompts can lead to the same safety failure. We conclude that this presents an issue for current AI development and for AI Safety, and leave a more comprehensive understanding of this phenomenon as an important challenge for future work.