Probing sycophantic follow-through across post-training stages with forced response openers
Sonnet Xu ⋅ Roxana Daneshjou
Abstract
To better understand sycophantic behaviors, this work tries to disambiguate how these behaviors occur in response initiation versus follow through. By prefixing model generations with forced openers, ranging from neutral commitment to explicit agreement, we evaluate how post-training influences sycophantic continuations. Applying this protocol to OLMo 3 7B checkpoints across base, supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning with verifiable rewards (RLVR) stages exposes a stark behavioral divergence between SFT and DPO that standard unforced generation masks. Given an acknowledgment opener, 62 of 154 paired items shift from capitulation under SFT to non-capitulation under DPO, with zero reversals (exact $p = 4.3 \times 10^{-19}$). However, over these same items, the length-normalized log-probability gap favoring wrong over correct candidates increases by 0.49. Medical questions exhibit a similar behavioral shift but an opposing contrast shift at deeper forced-agreement depths. Steering generations via a mean SFT-to-DPO residual-stream offset partially reproduces the behavioral effect on held-out items across domains. Controls reveal, however, that this offset is prompt-generic and lacks reliable cross-domain transfer. Forced openers therefore uncover post-training differences in sycophantic follow-through obscured by standard evaluations, while distributional and mechanistic interventions indicate these dynamics stem from domain-dependent mechanisms rather than a single unifying structure.
Chat is not available.
Successful Page Load