Steering Both Ways: Bidirectional Conditional Safety Steering in a Learned Contrastive Safety Space
Abstract
As language models are deployed in increasingly consequential settings, safety alignment must balance two demands: refusing harmful requests while remaining helpful on benign ones. However, safety should be judged by whether a model refuses the right requests, not by how often it refuses. Simply increasing refusal pressure can suppress harmful responses while making benign, sensitive-looking requests unanswerable, trading one failure for another. We present D3-CSS, an inference-time framework that separates harm detection from refusal execution. A shared two-task supervised-contrastive projection learns the safety geometry, with gradient-conflict resolution and row-orthonormal coordinates. A cached raw-state harm score and a running learned-space refusal score then drive opposing soft gates: D3-CSS restores refusal for harmful compliance, withdraws it for benign over-refusal, and suppresses edits when detection and behavior agree. We compare with seven external defense configurations and a probe-gated control, selecting each tunable method per model on development data. We summarize harmful-response and over-refusal rates by their macro-average, bidirectional safety error (BSE). Across seven safety sources and three task-utility benchmarks, D3-CSS is the only evaluated method that lowers both rates on both main models and significantly reduces BSE on each. Relative to the unsteered bases, BSE falls by 33.3% on Llama-3.1-8B and 38.5% on Gemma-2-9B, the largest reductions among the evaluated methods. General-purpose utility is largely preserved. A probe-gated control and development-model ablations support the roles of conditional routing and learned geometry, while a separate Qwen2.5-3B evaluation shows the same two-axis improvement.