Speculative Self-Distillation enables Efficient Knowledge Internalization
Abstract
Many practical deployments require language models to internalize knowledge that was absent from pretraining, such as proprietary corpora, continually updated facts, or specialized domain skills. Self-distillation has emerged as a natural approach for this task. A model conditioned on the source material acts as the teacher, and an unconditioned student is trained to reproduce the teacher's behavior closed-book. We propose Speculative Self-Distillation (SSD), an efficient per-token mixed-policy method for self-distillation. The student generates each rollout by default, while the teacher intervenes only at positions where its next-token distribution substantially diverges from the student's. This design keeps training rollouts close to the student's test-time behavior, unlike off-policy distillation with teacher-generated prefixes, but avoids spending teacher compute on the many low-signal tokens encountered by fully on-policy methods. In effect, SSD concentrates supervision precisely where privileged context changes the prediction. On three closed-book QA benchmarks measuring fact recall, knowledge updates, and skill acquisition, SSD matches on-policy distillation with 45\% fewer supervised tokens on average and consistently outperforms off-policy distillation on the held-out test set. SSD therefore establishes a new accuracy-efficiency Pareto frontier for self-distillation. Code is available at https://anonymous.4open.science/r/Speculative-Self-Distillation/.