Strong Teacher Not Needed? On Distillation in LLM Pretraining
Abstract
Knowledge distillation generally assumes a strong-to-weak relationship where stronger teachers yield better students. In this work, we examine this assumption about distillation in large language model (LLM) pretraining. By varying architecture sizes and training token budgets, we create strong-to-weak, same-level, and weak-to-strong teacher-student relationships. We study distillation's effectiveness under these relationships with different mixes between language and distillation loss. Three findings emerge: (1) with proper loss mixing, weak-to-strong and same-level distillation improves over standard pretraining, where even small and undertrained teachers benefit large students; (2) making the teacher stronger can lead to saturated or even reversed gains; (3) distillation improves generalization (out-of-domain, downstream) more readily than in-domain fitting. Our results provide practical guidance for choosing the teacher model with the proper loss for more effective LLM pretraining distillation.