Knowledge Distillation for Efficient Multilingual Speech Recognition: A Cross-Lingual Case Study on Swahili, Yoruba, and Hausa
Abstract
Deploying speech-enabled systems in low-connectivity, resource-constrained settings requires models that are small and fast, not just accurate. We study whether knowledge distillation (KD) can compress a multilingual speech recognition model while preserving accuracy on three under-resourced African languages: Swahili, Yoruba, and Hausa. Using a frozen \textit{Whisper-small} teacher and a \textit{Whisper-tiny} student, we compare a student trained with standard cross-entropy against a student trained with an additional KD loss, under an identical compute and data budget. The distilled student outperforms the size- and compute-matched baseline in all three languages, and also outperforms a Distil-Whisper-style pseudo-labeling baseline built on the same teacher. We assess robustness with bootstrap confidence intervals and three-seed replication: the Hausa gain is statistically distinguishable from test-set noise on its own, while the smaller Swahili and Yoruba gains are not individually significant at our test-set size, though both are directionally consistent across every one of nine independent training seeds. The resulting student uses 84\% fewer parameters and 84\% less memory than the teacher, with 59--68\% lower inference latency. We report these results, including where our evidence is and is not conclusive, and discuss implications for offline, on-device deployment in low-connectivity regions.