VoluCore: Spanning Teacher Representations with Volumetric Coresets for Data-Efficient LLM Distillation
Wang Xi ⋅ Yue Wang
Abstract
Data-efficient distillation of large language models depends not only on the student architecture, but also on which teacher examples are selected for training. Existing selection strategies often treat examples as independent or density-weighted samples, which can over-represent redundant high-frequency patterns while under-covering geometrically distinct regions of the teacher representation space. We study a volumetric alternative for distillation data selection. Our method, VoluCore, selects compact coresets by maximizing the regularized log-determinant of the Gram matrix formed by normalized teacher features. For any fixed regularization parameter, this objective is monotone submodular, giving a standard greedy $(1-1/e)$-approximation guarantee. To make the criterion practical at LLM scale, we implement greedy selection with efficient Cholesky/Gram-Schmidt rank-one updates after feature extraction. Across math, code, and instruction-following distillation settings, VoluCore reaches high-recovery performance with fewer selected examples than random, uncertainty-based, and distance-based selection baselines, while remaining competitive with gradient-based selection at substantially lower end-to-end cost. Additional spectral, layer, and task-coverage analyses indicate that VoluCore improves representation-space coverage without sacrificing common high-density capabilities. These results support log-determinant coverage of teacher representations as a simple, scalable, and theoretically grounded criterion for constructing compact distillation sets.
Chat is not available.
Successful Page Load