Equivalent Text, Unequal Predictions: Auditing Unicode Robustness in African News Headline Classification
Abstract
Affordable language technology must remain reliable under the representations its users supply. We distinguish canonical Unicode changes, which preserve text information, from diacritic deletion, which may destroy it. We audit six TF-IDF/linear-SVM pipelines on 2,608 held-out MasakhaNEWS headlines spanning Yoruba, Igbo, Hausa, Swahili, Nigerian Pidgin and French, with full, 64- and 128-code-point inputs. For the same baseline model, converting composed NFC text to canonically equivalent decomposed NFD text reduces full-headline macro-F1 from 83.17 to 59.06 in Yoruba and from 73.40 to 59.28 in Igbo. Paired bootstrap 95% intervals for these drops are [18.89, 29.55] and [7.64, 19.90] percentage points. French changes 11.90% of predictions despite a net F1 drop of only 0.45 points, illustrating that average quality can conceal representation instability. NFC preprocessing guarantees canonical invariance but does not repair deleted marks. For Yoruba and Igbo, augmentation with mark-removed training views improves performance on stripped text over a duplicated-clean control; its advantage over simple stripping remains uncertain. We also measure CPU throughput and serialized model size to characterize the computational costs of these remedies. This study provides an inexpensive evaluation protocol that separates information-preserving invariance, information-loss sensitivity and computational cost. Results are limited to one headline-classification dataset and synthetic perturbations; validation with speakers, additional domains and target devices remains future work.