On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
Abstract
Recent work has identified rollout policy as a potentially central distinction between post-training methods, arguing that on-policy learning can improve generalisation, reduce catastrophic forgetting, and produce sparser parameter updates. However, comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. We study this question in a controlled strong-to-weak distillation setting, matching the teacher, supervision signal, optimisation procedure, and KL direction while varying whether training trajectories are generated by the student or teacher. Across two model families and reasoning tasks spanning medical, scientific, and arithmetic domains, we find no consistent advantage for on-policy distillation (OnPD) over off-policy distillation (OffPD) in final performance, forgetting, update sparsity, or generalisation. Instead, learning rate primarily determines forgetting and update sparsity, while KL direction more strongly shapes generalisation, output coverage, and performance after subsequent reinforcement learning with verifiable rewards (RLVR). Rollout policy still matters in specific regimes: OnPD becomes less stable without gradient clipping, while OnPD with reverse KL is more robust to learning incidental teacher style. These results suggest that, within the controlled distillation regime studied here, rollout policy alone is insufficient to explain many behaviours commonly associated with on-policy post-training, and that the cheaper OffPD regime should serve as a standard baseline when evaluating OnPD.