OASIS: Online Adaptive Steering for In-Training Safety of LLMs
Abstract
Fine-tuning is essential for adapting Large Language Models to downstream tasks. However, this process can also inadvertently erode the critical safety alignment even when the fine-tuning data appears benign. Prior methods either introduce safety regularizers that may conflict with the primary training objective, or rely on fixed mechanisms that fail to adapt to fine-tuning dynamics. To remedy this, we reframe safety preservation as a training-time intervention and propose Online Adaptive Steering for In-Training Safety (OASIS), enabling continuous online recalibration of the steering direction during fine-tuning. Specifically, OASIS tracks a misalignment direction in activation space, then applies example-adaptive activation steering to absorb misalignment-inducing updates during training. Across three misalignment behaviors and three fine-tuning data regimes, including fully misaligned, mixed, and benign data, OASIS consistently improves safety robustness without sacrificing downstream performance.