Implicit Robustness Reweighting in Private Fine-Tuning
Abstract
Modern ML systems may require several forms of protection simultaneously. However, these protections can interact during training. We study this problem at the intersection of differential privacy and adversarial robustness. We compare strategies that bound the gradients of a composite robustness objective jointly or separately, while holding the objective, sensitivity bound, and noise calibration fixed. Through a geometric analysis, we characterize how these strategies preserve or alter the relative contributions of task and robustness terms. In particular, separate bounding can induce effective robustness weights that differ across records and change during training. Experiments with TRADES show higher adversarial robustness with joint bounding under these matched conditions. Measurements on public examples reveal substantial variation in effective weights even when their median is close to the specified loss weight. Together, these findings show that the interaction between a robustness objective and private optimization must be understood through the resulting training updates, not the loss alone.