Beyond Answering: Auditing Multilingual Public Service Triage
Abstract
Public-service assistants must decide when to answer, ask for a missing fact, or refer a case to an official route. We audit an earlier GlobalSouthML study’s frozen 90-item Sinhala, Tamil, and Swahili benchmark and five systems. New analyses separate routing from clarification-slot errors, evaluate 30 Sinhala–Tamil pairs, and quantify service-cluster uncertainty. Qwen 2.5 1.5B asks on every case but selects the correct slot on only 3 of 30 clarification cases. Qwen 2.5 3B answers 55 of 60 cases requiring clarification or escalation. Cross-language agreement is also misleading: the 1.5B model agrees on all 30 pairs, yet resolves neither lan- guage correctly on 28. Sonnet 4.6 achieves 76/90 joint action-and-slot correctness and resolves both versions of 22/30 pairs. These are new analyses of existing predictions, not new model runs. The findings motivate evaluation that combines routing, clarification, and paired correctness before public-service deployment.