UniSD: Towards A Unified Self-Distillation Framework for Large Language Models
Abstract
Self-distillation (SD) offers a promising path for adapting large language models (LLMs) without relying on stronger external teachers. However, self-distillation in autoregressive LLMs remains challenging because self-generated trajectories are free-form, correctness is task-dependent, and plausible rationales can still provide unstable or unreliable supervision. Existing methods mainly examine isolated design choices, leaving open when self-distillation helps, which controls make it reliable, and how different mechanisms interact. We introduce UniSD, a Unified framework for systematically studying Self-Distillation. UniSD formulates self-distillation as reliability-aware on-policy distillation: self-derived teacher views score student-sampled trajectories, and the resulting distillation objectives update the model parameters. For controlled comparison, UniSD groups the mechanisms by three functional roles: supervision reliability, representation alignment, and training stability. Across six benchmarks and six models from three model families, UniSD yields several findings. Static imitation yields limited and task-dependent gains, while the on-policy act as stronger teachers. Agreement trades peak performance against robustness and varies across context construction. Guided by these insights, we construct UniSDfull, an integrated pipeline that achieves the strongest overall performance, improving over the base model (+5.4) and the strongest baseline (+2.8). Our code is available at https://anonymous.4open.science/r/UniSD