You Don’t Need Aligned Representations: Knowledge Distillation via Random Prototype Spaces
Hamza Etcibasi ⋅ Ramazan Gokberk Cinbis
Abstract
A widely held assumption in knowledge distillation is that teacher and student representations must occupy a compatible or explicitly aligned feature space. We propose RPKD (Random Prototype Knowledge Distillation), which discards this assumption by projecting both networks into a shared set of randomly initialized, frozen prototype vectors, requiring no architecture-specific adapters, no class-specific alignment, and minimal architectural assumptions. Two branches handle the transfer: a logit branch aligning global prototype similarity distributions and a feature branch matching spatially-resolved prototype activation maps. Both operate within the same fixed prototype space, which can be interpreted as a random feature embedding that approximately preserves inter-sample similarity structure, keeping RPKD agnostic to the internal dimensions and inductive biases of either network. In low-category settings, decoupling the prototype vocabulary from the task label space yields richer supervision than the label space alone, an advantage absent from logit-based methods, whose supervisory signal collapses as class count shrinks. This advantage is especially relevant in real-world scenarios where the label space is typically constrained. Extensive experiments demonstrate the effectiveness of RPKD, achieving a maximum gain of $\mathbf{+8.26\%}$ over OFA on CIFAR-100, $\mathbf{+12.46\%}$ on ImageNet-100, and $\mathbf{+1.34\%}$ on ImageNet-1K. Source code will be released upon acceptance. Source code will be released upon acceptance.
Chat is not available.
Successful Page Load