Sequence-Level Knowledge Distillation Collapses Uncertainty and Erodes Hallucination Detection
Abstract
Sequence-level knowledge distillation (SeqKD) is widely used to transfer the capabilities of large language models to smaller students, but the teacher responses used as training targets can themselves contain hallucinations. We study what happens to teacher-generated hallucinations and their associated uncertainty as distillation proceeds. Standard SeqKD retains only one or a few sampled responses per input. It records what the teacher answered but not the distribution from which the answer was drawn. Theoretical analysis shows that continued learning on a single target drives the student towards a concentrated output distribution, regardless of target correctness. This predicts hallucination inheritance, uncertainty collapse, and a progressive loss of uncertainty-based hallucination detectability. Experiments across four teacher–student model pairs on SimpleQA confirm these predictions. Students increasingly reproduce hallucinated teacher targets while becoming more certain of them, and the AUROC of semantic entropy for detecting hallucinations falls from teacher values between 0.73 and 0.83 to near chance at later student checkpoints. We then evaluate three interventions that use teacher uncertainty during corpus construction. Filtering high-uncertainty examples slows detectability loss, while replacing their targets with abstentions improves selective answering. Sampling multiple teacher responses preserves detectability throughout training, but the cost grows with the number of retained responses. None of the three interventions is best in every respect, and maintaining hallucination detectability through distillation remains an open problem.