Adapting Multilingual Self-Supervised Speech Models for Low-Resource African Child Speech Recognition
Abstract
Despite major advances in automatic speech recognition (ASR), the recognition of children’s speech remains significantly less accurate than adult ASR. ASR systems trained primarily on adult speech perform poorly when evaluated on child speech because of the substantial acoustic mismatch between adult and child speech. Although multilingual self-supervised speech models such as XLSR-53 and mHuBERT-147 have shown strong multilingual ASR performance, their adaptation to low-resource African child speech remains largely unexplored. This paper investigates the adaptation of multilingual self-supervised speech models for African child speech recognition and proposes an ASR Adaptation Framework comprising four stages: Data Adaptation, Representation Adaptation, Model Adaptation, and Learning Adaptation. Preliminary experiments evaluate adult-trained models on adult and child speech, investigate adult speech augmentation using child-like acoustic transformations, and fine-tune the selected model using limited child speech. Results confirm a substantial adult–child acoustic mismatch, with XLSR-53 providing the strongest child-speech baseline. Adult speech augmentation improved adult WER from 20.2% to 17.6%, but the improvement did not transfer to child speech. Fine-tuning on limited child speech reduced child WER from 96.5% to 96.2%, while adult WER increased to 28.4%. These findings show that adult-only augmentation is insufficient to bridge the adult–child acoustic mismatch and that limited African child speech data remains a key bottleneck. The proposed framework provides a foundation for further investigation of child-specific augmentation, voice conversion, and parameter-efficient fine-tuning for low-resource African child speech recognition.