Multilingual Health and related Nutrition Question Answering for Low-Resource African Languages
Abstract
Speakers of low-resource African languages risk exclusion from digital health and nutrition information due to limited language support in existing platforms. We present early-stage results toward a voice-enabled, multilingual question-answering (QA) system that enables speakers of Kinyarwanda, Khelobedu, and Kidawida to ask health and nutrition questions in their own language and receive grounded responses, with automatic speech recognition (ASR) and translation results re- ported here for Kinyarwanda and Kidawida. Our framework integrates ASR, a knowledge graph, and large language models (LLMs) for response generation and cross-lingual translation. For ASR, Conformer-CTC, RNN-Transducer (RNNT), and Whisper were fine-tuned on a balanced bilingual set (27.5 hours/language) and compressed using structured pruning and quantization. The optimal strategy was 20% pruning with int8 quantization for Conformer-CTC and RNNT, reducing model size by 68.8% and 70.2%, respectively, and int8 quantization alone for Whisper, reducing size by 70.2%. Following compression, Conformer-CTC WER increased from 11.10% to 12.0% on Kinyarwanda and 26.48% to 29.2% on Kidaw- ida, while RNNT WER increased from 10.76% to 12.1% and 17.57% to 18.6%, respectively. Whisper achieved compressed WERs of 9.4% and 11.6%, but showed the largest validation-to-test degradation on Kidawida. Conformer-CTC offered the best size-accuracy-efficiency balance for deployment. For Kinyarwanda-English translation, three multilingual architectures were fine-tuned on 226,333 sentence pairs and evaluated on an 11,318-pair held-out test set spanning general and health domains. On the health subset, pretrained NLLB-600M achieved 1.30 BLEU/13.47 chrF (EN→RW) and 7.13/28.70 (RW→EN) before task-specific fine-tuning; after fine-tuning, NLLB-600M achieved 52.65 BLEU/73.37 chrF. NLLB-600M was then adapted to Kidawida via LoRA using 18,199 Kidawida sentences mixed with equal Kinyarwanda data. On a 5,073-pair Kidawida test set, adaptation improved per- formance from 0.16 BLEU/13.45 chrF to 10.76/38.86 (EN→Kidawida) and from 1.68/13.40 to 28.61/47.21 (Kidawida→EN), while Kinyarwanda BLEU decreased by at most 1.36 points. The knowledge graph draws on four primary documents labelled gold and 19 secondary documents labelled silver, covering hypertension, type 2 diabetes, osteoarthritis, and autoimmune conditions. The secondary docu ments were chunked and then filtered using a calibrated cosine similarity threshold of 0.4 against the primary documents. The resulting graph consists of 8 classes and 13 object properties. Entities are grounded to 4 biomedical and food ontologies, with indigenous foods represented through a local namespace; citation validation removes triples whose cited text cannot be verified in the source. These module- level results provide early evidence toward a multilingual, voice-enabled health QA system for low-resource African languages. End-to-end integration, quantitative reporting of the knowledge graph evaluation, and extension to Khelobedu remain ongoing work.