Safety Erosion from Benign Narrow Fine-Tuning
Abstract
Companies are fine-tuning models on their own customer conversations, and we have little idea what that does to safety. Fine-tuning GPT-4.1, Qwen3.5-27B and Llama-3.3-70B on AI-drafted, human-edited replies to one retailer's real customer service emails moves safety-relevant behavior on benchmarks unrelated to customer service, in every fine-tune we ran. Harm avoidance, deception, sycophantic mirroring, and value choices all shift, general capability holds, and no adversarial content appears anywhere in the training data. Model families do not act the same: harm avoidance falls sharply on two families and not the third, while all three lie more about who made them. The erosion also does not track the content. We trained three variants of the same emails, flattening the tone, deleting the sales tactics, and transposing the product, and each changed little. Of our two controls, a personality dataset teaching low openness erodes as much as ours or more, while a self-distilled dataset is the only arm that stays at or near baseline. Retraining identical data with a different seed moves the result further than any change we made to the data on some benchmarks.