Efficient On-Device Language Models for Low-Resource African Languages: A Pilot Study of Quantization Effects on Swahili Task Performance
Abstract
Efficiency research on large language models (LLMs) is overwhelmingly bench- marked on English, while multilingual NLP research on low-resource African languages like Swahili rarely accounts for the deployment constraints of real-world edge and offline settings. This gap matters directly for the Global South: cheap, offline-capable AI assistants could meaningfully serve communities with limited connectivity and hardware, but only if compressed models retain language com- petence. We present a pilot study measuring the accuracy-efficiency tradeoff of quantizing three open small language models (Llama 3.2 3B, Qwen 2.5 3B, Gemma 2 2B) at three precision levels (fp16, int8, int4) on a 40-item Swahili benchmark spanning factual question answering, proverb interpretation, code-switching ro- bustness, and translation. To overcome dataset constraints in low-resource settings, benchmark items and candidate reference answers were initially drafted with LLM assistance (Claude) and subsequently underwent 100% review by a native Swahili speaker (the author), who corrected dialectal variations, cultural nuances, and proverb phrasing to ensure full target accuracy, scoring each item on a 0–2 correct- ness scale. We find two distinct patterns rather than one blended effect: aggregate task accuracy varies by quantization level and by task; translation was weakest overall and degraded most under int8 (0.33 → 0.17 on a 0–2 scale), while code- switching was comparatively robust (0.50 → 0.67) with Llama 3.2 3B emerging as the strongest and most compression-robust model overall (0.80 at both fp16 and int4). Meanwhile, three qualitative failure modes — prompt-echoing, confident fabrication, and degenerate repetition — occur at comparably high rates already at full precision and do not meaningfully increase under compression (fp16/int8/int4 respectively: confident fabrication 30.0%/26.7%/26.7%; degenerate repetition ex- actly 20.0% at every level; prompt-echoing 13.3%/13.3%/16.7%, all as a share of scored responses). Every one of the nine model-quantization configurations fabricated a specific, non-existent local business when asked for a recommendation, and models varied widely in how often they responded in English despite Swahili prompts (14.2%–67.5% across models). On a single NVIDIA T4 GPU (Kaggle), int8 quantization unexpectedly increased mean per-response latency roughly 8-fold over fp16 (5.3s → 45.5s) despite a smaller memory footprint (e.g., Llama 3.2 3B: 6.42 GB at fp16 vs. 1.60 GB at int4), showing that memory savings and inference speed do not necessarily co-optimize on this hardware. Given the small per-cell sample size (n = 10), we report these as descriptive pilot-stage trends rather than statistically tested effects, and see this work as motivating larger, adequately powered follow-up studies across more African languages and models. We release our benchmark and evaluation pipeline to support this direction.