Dual-Null LoRA: Decoupling Fine-Tuning Updates from Safety Alignment
Abstract
The safety alignment of large language models, typically established as part of post-training pipelines, is known to degrade under subsequent fine-tuning. This phenomenon occurs not only when fine-tuning on deliberately harmful examples but also using harmless, domain-specific utility data. In a first-order approximation of fine-tuning dynamics, one fine-tuning step on a potentially harmful training prompt changes the probability of responding in an aligned way to an observed prompt through the empirical neural tangent kernel that couples training and observation prompt. We introduce dual-null LoRA, a standard LoRA adapter built to eliminate this coupling: Our adapter is initialized in approximate null spaces of both, activations and gradients, of a small safety calibration set, and every update is projected back into these null spaces, yielding, up to calibration error, an alignment change of zero to first order for any fine-tuning example and objective. Evaluated on two base models (OLMo-2-1B, SmolLM2-1.7B) and three fine-tuning datasets (SST-2, AG News, GSM8K), dual-null LoRA stays within 1.8 AdvBench harmfulness (HS) and 0.8 attack-success (ASR) points of the pre-fine-tuning baselines, even at 20% harmful examples (undefended LoRA: 59.8 and 66.2).