A critical look at self-distillation for reasoning
Abstract
In self-distillation, a teacher is constructed by prompting a student model with privileged information or environmental feedback, and logit-level feedback is derived from the difference between the two models' output distributions. This paradigm has recently gained traction as an alternative to policy-gradient reinforcement learning, promising dense supervision where reward signals are sparse. This paper critically evaluates whether self-distillation methods indeed improve performance in reasoning tasks using privileged information. We report three findings: (1) Self-Distillation Policy Optimization only improves performance on LiveCodeBench v6 problems seen during training, and its reported gains in science reasoning and tool use are largely recoverable through supervised fine-tuning; (2) the gains of On-Policy Self-Distillation in math can be attributed to distilling thinking-mode behavior into the non-thinking student, a mechanism that does not require privileged information; (3) student-teacher self-distillation objectives rely on final answer correctness to discriminate between correct and incorrect student solutions, providing little signal about reasoning correctness. Together, these results indicate a need for robust methods that leverage privileged information and feedback.