Steering Vectors as a Training Signal in LLM Post-Training
Tiejin Chen ⋅ Maunil R Vyas ⋅ Huaiyuan Yao ⋅ Hua Wei
Abstract
Steering vectors (SVs) are residual-stream directions that shift the behavior of LLMs when added to hidden states at inference time. While prior work has mainly used SVs as post-hoc controls, we ask whether the same directions can also guide parameter updates during post-training. We study this question across supervised learning, including SFT and distillation, and reinforcement learning with verifiable rewards (RLVR). Our framework considers two intervention regimes: a hidden-state hook during teacher-forced supervised training, and an advantage-repair mechanism for zero-variance groups in GRPO. For each regime, we analyze three design axes, which are the steering direction, its placement, and the operation used to combine it with training. In supervised settings, SV injection improves over plain SFT and distillation, with gains scaling with the off-policy gap and reaching up to a $12\%$ improvement when distilling from thinking-mode teachers. In RL settings, applying a correctness-contrast direction as advantage repair on all-wrong groups recovers a learning signal that transfers to held-out math benchmarks. We further identify three failure boundaries. In detail, code generation can break the direction, cross-family layer transfer can break the placement, and inference-time application on an SV-trained checkpoint can break the intervention regime. Together, these results establish steering vectors as a conditional but effective training signal for LLM post-training. Our code can be found in Appendix Section A.
Chat is not available.
Successful Page Load