Reward-Aligned Distillation: Distill Your Model towards Higher Reward
Abstract
Reinforcement learning with verifiable rewards improves the reasoning capabilities of large language models, but remains computationally expensive and sample inefficient. On-policy distillation (OPD) provides dense token-level supervision from a stronger teacher, yet standard OPD distills every token regardless of whether the teacher's preference actually improves task reward. Existing token-selection methods rely on proxies such as uncertainty or teacher--student divergence, which need not identify reward-improving supervision. We introduce Reward-Aligned Distillation (RAD), a bilevel framework that directly selects teacher signals based on their predicted effect on reward. RAD derives a token-level lookahead score that measures whether distilling a token is expected to improve the reward of the updated student, and masks tokens with detrimental supervision. Under a rollout-matched budget, RAD improves mean accuracy on four competition-math benchmarks over the strongest baselines by 8.0 percentage points for Qwen3-1.7B and by 4.4 points for Qwen3-4B-Instruct-2507. RAD also outperforms all baselines on LiveCodeBench for both models, showing that explicitly aligning token-level distillation with task reward substantially improves OPD.