Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
Abstract
Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student without it, both from the same model. This guidance is valuable when the student has failed. On successful rollouts, however, the same mechanism backfires: it overwrites the choices that produced the success, suppressing the student's own reasoning. We therefore propose reading the self-distillation signal in reverse: when the student succeeds along a path the teacher would not have predicted, these tokens reflect its self-driven reasoning, and reinforcing them yields valuable exploration grounded in the student's own success. Building on this, we propose RLRT (RLVR with Reversed Teacher), which augments GRPO by amplifying these tokens on correct rollouts. Across base, instruction-tuned, and thinking-tuned Qwen3 checkpoints, RLRT substantially outperforms self-distillation and exploration-based baselines. Its gains are largest on base models, whose output distributions prior post-training has not yet narrowed. Our analysis shows that RLRT reorganizes rather than merely sharpens the pretrained policy's reasoning distribution, shedding light on how post-training reshapes, and depends on, what pre-training leaves behind.