ASR for Kiswahili Early-Grade Read-Aloud Assessments in Tanzania
Abstract
Roughly 70\% of 10-year-olds in low- and middle-income countries cannot read and understand a simple text, but the oral assessments that diagnose this are costly and slow, requiring one-on-one administration. Automatic speech recognition (ASR) offers a route to scale, yet assessment needs something ordinary ASR is not built to do: record what a child actually said, mispronunciations, hesitations and self-corrections included, rather than what they meant. Child speech data for low-resource languages is also scarce. We present an ASR system for Kiswahili early-grade read-aloud assessment in Tanzania: a 116M-parameter FastConformer encoder with a CTC BPE decoder, fine-tuned on Kiswahili adult speech and then on real child speech from EGRA reading tasks. We pair it with a reproducible benchmark evaluating systems both orthographically (word error rate) and in IPA (phoneme error rate), alongside a miscue error rate and per-task agreement with human scorers. Across 19 models and 25 configurations, our model takes the three highest positions in both views, reaching 18.30\% WER and 6.81\% PER against 45.30\% and 18.01\% for the strongest alternative evaluated. Errors concentrate on isolated letters and syllables rather than connected passages, underlining the need to evaluate on authentic early-reading tasks.