On the Off-Policy Teacher in On-Policy Distillation
Abstract
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm that combines student-on-policy data collection with dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. Empirically, teacher model achieves lower performance when continuing from student prefixes. This consequently degrades the quality of supervision provided to the student. The problem becomes particularly severe in long reasoning trajectories, where the discrepancies accumulate over reasoning steps. To address this issue, we propose Student-Conditioned On-Policy Updates for Teachers (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using GRPO: the teacher generates continuations from student prefixes and learns from verifiable outcome rewards. Through controlled experiments, we empirically demonstrate the existence and progression of teacher-side off-policy degradation and show that SCOUT effectively mitigates it. Experiments across multiple teacher--student configurations, model scales, and reasoning domains demonstrate that adapting the teacher to student-generated prefixes consistently improves the effectiveness of on-policy distillation.