Projected Distillation: On Learning to Solve Hard Problems from Close, Correct Targets
Abstract
On-policy learning is a cornerstone of post-training: by training a model on its own generations, it limits distribution shift and helps preserve existing capabilities. On hard problems, however, this signal thins out. On-policy RL provides no learning signal when every attempt fails and rewards are uniformly zero. On-policy distillation, which supervises the student's erroneous rollout token by token towards a teacher, is inefficient because feedback on later tokens remains conditioned on a mistaken prefix. Correct behavior can instead be imported off-policy from a gold or teacher solution, but such targets sit far from the model's generations and erode existing capabilities. We therefore view on-policyness as a continuum and ask where along it supervision should lie. We propose projected distillation, which trains on the correct solution closest to the model's failed attempt, a minimal correction refreshed online as the model evolves. Across two students (Qwen3-8B, OLMo-2-7B) and five tasks spanning math, code, science, and tool use, projected distillation achieves the best specialization--retention frontier, outperforming the top baselines. We show that this trade-off is governed by how far the supervision sits from the model's own distribution, so minimally editing its own wrong attempt buys new capability at almost no cost to old ones. Our ablations identify the key ingredients for near-on-policy external supervision: correction size, target-solution selection, correction refresh schedule, training objective and correction source. Theoretically, we formalize the continuum: near-on-policy corrections yield DAgger-like linear regret, while fully off-policy targets recover the quadratic regret of behavior cloning.