The Reach of Emergent Misalignment: How Scale and Similarity Influence Cross-Domain Transfer
Alexander Arutchev ⋅ Atharva Bhargude
Abstract
Emergent Misalignment (EM) occurs when harmful fine-tuning in one domain degrades a model's alignment in unrelated domains, complicating safety evaluation for deployment-specific fine-tunes. This study presents the largest cross-domain transfer study of EM, fine-tuning 17 checkpoints spanning five families (Qwen2.5-Coder 0.5B--32B, Gemma-3 1B--27B, Llama-3 1B--8B, OLMo-2 1B--13B, and Phi-3 3.8B--14B) on each of 11 harmful domains and evaluating every fine-tuned model on 11 domains, using a GPT-4o-mini judge validated with human ratings. EM is prevalent across domains: fine-tuning on almost any one domain raises misalignment on most others at every model size tested. Scale changes how EM appears rather than reducing measured misalignment. Coherence rises steeply with size, making larger models more fluent about misaligned responses. Domain similarity becomes a more reliable predictor of transfer risk as size increases across all five families, reaching $r = -0.616$ ($p < 0.001$) for Gemma-3 27B and explaining up to 38\% of the variance. A benign secure-code control supports an EM interpretation over generic catastrophic forgetting. Forced reasoning roughly halves EM in Qwen3 at 4B and above, with little change below.
Chat is not available.
Successful Page Load