MinSteer: Minimal-Pair Steering via Two-Stage Cached Continuation
Abstract
Do behavioral styles in large language models correspond to reusable directions in residual space, or do they mainly appear after averaging many noisy contrastive examples? We introduce MinSteer, a two-stage cached-continuation method for estimating a contrastive residual direction (CRD) within one decoding trajectory. Stage-X generates a continuation in one style; a style-flip suffix is then processed using Stage-X's final KV cache; Stage-Y continues from this inherited trajectory. The CRD is computed from generated-token residuals only, excluding prompt and suffix tokens from the average. Across sentiment transfer, politeness control, and a ParaDetox-based bidirectional diagnostic, MinSteer extracts effective steering directions from very small contrasts. On TweetEval sentiment transfer, one MinSteer contrast outperforms a CAA baseline averaged over 50 independent pairs. Japanese keigo CRDs transfer to English and Chinese outputs and outperform trajectory-free TwoSeq controls. In ParaDetox, MinSteer produces bidirectional Detoxify movement under our rewrite setup, while exposing semantic-preservation trade-offs. The directions transfer across tested scenarios and languages and show clear depth dependence. Together, these results show that cached self-rewrite trajectories provide a cleaner signal for residual-space behavior directions than independent-pair averaging.