Learning from Negative Policy: Revisiting Rollouts in On-Policy Distillation
Abstract
On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. While recent studies have extensively explored how to improve the distillation signal of OPD, the role of the rollout policy itself remains less understood. In this work, we study negative-policy distillation under the hypothesis that rollouts from a lower-capability model, referred to as the negative-policy model, can serve as a negative reference, exposing behaviors that the student can learn to suppress rather than imitate. Interestingly, training with negative-policy rollouts consistently outperforms standard on-policy rollouts across model scales, generation modes, reasoning tasks, and various OPD formulations. Our analyses further show that the model trained with negative-policy distillation assigns lower probability to tokens favored by the negative-policy model, supporting our hypothesis that negative-policy rollouts serve as a negative reference similar to rejected responses in preference optimization. Furthermore, negative-policy distillation maintains consistent gains as the number of training steps increases, whereas standard OPD becomes less effective, providing new insight into the role of rollout policy in OPD.