Measuring the Safety Cost of Preventative Persona Steering
Abstract
High-level behaviors such as evil, sycophancy, and hallucination have been shown to correspond to directions in a model's activation space \citep{chen2025persona}. These directions, known as persona vectors, let practitioners suppress undesirable traits like sycophancy and hallucination in two ways: inference-time suppression (subtracting the vector at generation) and preventative steering (adding the vector during fine-tuning, so the optimizer never needs to encode the trait in weights). \citet{chenarditi2025followup} recommend preventative steering as the better option, reporting that it preserves MMLU where inference-time steering degrades it. Every validation of this claim to date measures capability alone. While preventative steering has been promoted as preserving general task performance, our preliminary audit reveals that it selectively alters refusal dynamics: static automated benchmarks show increased attack vulnerability, while standard safety prompts show sharpened refusal, highlighting that safety side effects are benchmark-dependent and cannot be inferred from capability scores alone.