The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
Abstract
Fine-tuning aligned language models on benign tasks (e.g. math tutoring) systematically breaks safety guardrails, even when training data contains no harmful content. While mechanistic approaches have shed light on where alignment resides in model weights, they lack the formal language needed to derive guarantees about when and why fine-tuning degrades it — leaving the field without principled tools for predicting or preventing alignment collapse. We develop such a framework through geometric analysis of parameter-space trajectories and apply it to understand the fragility of alignment in fine-tuning. While first-order analysis suggests orthogonal updates are safe, we prove this is illusory: the curvature of the fine-tuning loss induces second-order acceleration that systematically bends trajectories into alignment-sensitive regions. We formalize the central construct of our framework as the Alignment Instability Condition (AIC), three geometric properties that, when present, are sufficient to guarantee degradation. Our main result proves quartic onset of alignment degradation in training steps, determined by how sharply alignment depends on specific parameters and how strongly tasks couple to these parameters. These findings yield formal sufficient conditions under which gradient descent makes alignment inherently fragile, demonstrating the framework's capacity for theoretical guarantees. We further empirically validate the framework's foundations, showing that the Fisher Information Matrix governs the degree of safety degradation across diverse fine-tuning settings.