SURGE: Sparse Update Routing via Gated Entropy for Efficient Active Distillation
Tanvir Bhathal ⋅ Allen Shen ⋅ Gabriel Hübner ⋅ Dylan Zhou ⋅ Karthik Lakshmanan
Abstract
Post-training should correct a base model's failures without unnecessarily rewriting behaviors it already performs reliably. Yet standard on-policy distillation expends significant compute updating predictable tokens, while naive sparse masking destabilizes unselected behavior through shared parameters. We propose \textbf{SURGE}, a two-part active distillation framework that includes: (1) \textit{Dual-Gated Sparse Routing}, combining macro-level PRM verification with micro-level token-predictive entropy to identify complementary failure signals ($\le$ 2.90\% overlap), and (2) a \textit{Decoupled KL Objective}, which applies standardized teacher advantages to active failures while anchoring inactive positions to the reference policy. When distilling a Gemma4 31B teacher into a Gemma4 4B student across five benchmarks, SURGE consistently outperforms SFT, GRPO, PRM, and OPD. Relative to standard OPD, SURGE achieves accuracy gains across all benchmarks while cutting distillation FLOPs by an average of \textbf{50.9\%} (up to \textbf{74.9\%}) and reducing response verbosity by an average of \textbf{33.0\%} (up to \textbf{55.8\%}). The results show that routing and regularization not only determine improvements of targeted capabilities, but also which base-model behaviors are preserved.
Chat is not available.
Successful Page Load