Token Utility for On-Policy Distillation via Local Trust-Region Optimization
Yilan Chen ⋅ Vijay Lingam ⋅ Narayanan Sadagopan
Abstract
On-policy distillation (OPD) trains a student on its own rollouts while a stronger teacher provides dense token-level supervision. Recent work has shown that not all student-visited positions are equally useful for learning, but has characterized token importance through properties such as uncertainty or teacher--student disagreement. We study token utility through a complementary optimization-based view: rather than specifying a notion of importance a priori, we define a local optimization problem and let token utility emerge from its solution. Specifically, we ask how much the token-level distillation objective can decrease when each position is allowed the same local change in the student's predictive distribution. The Fisher geometry of KL yields the utility $s_t=g_t^\top F_t^\dagger g_t$, where $g_t$ is the logit gradient and $F_t$ is the categorical Fisher matrix. For reverse-KL distillation, this quantity reduces exactly to the student-weighted variance of the student--teacher log-probability ratio. We use this result to construct Fisher Utility Weighting (FUW), a practical mean-normalized weighting rule. In end-to-end distillation of Qwen3-1.7B-Base from Qwen3-8B, FUW reaches a Macro Avg@8 of $26.8$ and a pooled score of $39.1$ across seven mathematical reasoning benchmarks, the strongest aggregate performance among the evaluated student-training methods. Controlled studies further show that continuous utility weighting outperforms hard token selection using the same utility, indicating that the relative magnitudes of the utility scores can provide additional training signal. Finally, the same optimization framework extends to forward-KL distillation, where it induces the Pearson $\chi^2$ divergence as the corresponding token utility, demonstrating that the framework yields objective-specific notions of token utility.
Chat is not available.
Successful Page Load