Follow the Failure: Tracing Safety Degradation Back to Training Data
Abstract
Post-training a Large Language Model (LLM) on instruction data is known to sometimes degrade the safety behavior of the model, even when no training examples look harmful. We ask which examples in the supervised fine-tuning (SFT) mixture are responsible, and show that the degraded model itself supplies the signal needed to find them. We find that taking a fine-tuned model's own unsafe generations as an attribution target, embedding them and the training corpus by the gradients they induce, and removing the most aligned training examples before retraining from the base checkpoint reduces attack success rate on held-out HarmBench under transferred GCG suffixes from 9.4\% to 3.1\% and from 10.2\% to 3.8\%, whereas an LLM-as-a-judge baseline removing the same fraction attains only 8.8\% and 10.2\%, while improving performance on GSM8K or MATH 500. To make the attribution scalable, we develop an optimized gradient embedding method that restricts backpropagation to the final few transformer blocks. These findings suggest that filtering the first SFT mixture with signal drawn from post-training itself is a promising route to preserving safety.