The Reach of Emergent Misalignment: How Scale and Similarity Influence Cross-Domain Transfer
Alexander Arutchev ⋅ Atharva Bhargude
Abstract
Emergent Misalignment (EM) occurs when harmful fine-tuning in one domain degrades a model's alignment in unrelated domains, complicating safety evaluation for deployment-specific fine-tunes. This study presents the largest cross-domain transfer study of EM, fine-tuning 17 model checkpoints spanning five families (Qwen2.5-Coder 0.5B--32B, Gemma-3 1B--27B, Llama-3 1B--8B, OLMo-2 1B--13B, and Phi-3 3.8B--14B) on each of 11 harmful domains and evaluating every resulting fine-tuned model on all 11 domains, using a GPT-4o-mini judge validated with human ratings. EM is prevalent across domains: fine-tuning on almost any one domain raises misalignment on most others at every model size tested. Increasing model size changes how EM appears rather than showing a reduction in measured misalignment. Coherence climbs steeply with size, so larger models are more fluent about their misaligned responses. Semantic similarity between the domains becomes a more reliable predictor of relative transfer risk across all five families, reaching $r = -0.616$ ($p < 0.001$) for Gemma-3 27B and explaining up to 38\% of the variance. A benign secure-code control supports the interpretation that the effect is emergent misalignment rather than generic catastrophic forgetting. Finally, forcing a model to reason before it answers roughly halves fine-tuning-induced misalignment in Qwen3 at 4B and above, and changes little below that.
Chat is not available.
Successful Page Load