SAFE-DRIFT: Data Selection for Supervised Fine-tuning with Controllable Off-Target Drifts
Abstract
We propose SAFE-DRIFT, a data-selection framework for supervised fine-tuning that explicitly trades off target improvement against unwanted off-target drifts of model behavior. We formalize this objective as a constrained optimization problem: maximize gain on the target gradient subject to a budget on off-target drift, which we quantify by the Fisher information on the reference distribution. The resulting closed-form solution is a damped natural gradient with respect to the reference Fisher matrix. We further address the statistical estimation challenge in the high-dimensional setting, deriving a scalable algorithm that approximates the Fisher information and the target gradient in a shared low-dimensional subspace. We evaluate SAFE-DRIFT against state-of-the-art data selection methods across diverse domains including coding and medical question answering, yielding competitive target performance while reducing drift on specified off-target behaviors.