Scaling and Evaluation Across India: A High-Risk Pregnancy Chatbot for Community Health Workers
Sthitapragyan Parida ⋅ Saumitra Sharma ⋅ Sanskriti Midha ⋅ Jigar Doshi
Abstract
We report the design, deployment, and evaluation of a high-risk pregnancy chatbot that assists over $8{,}500$ community health workers (CHWs), a frontline cadre central to reducing maternal mortality, in India. Accessible through WhatsApp, a channel they already use, it integrates Large Language Models (LLMs) into their existing workflow for real-time support in speech and text across four Indic languages and their code-mixed variants. The paper delivers two things: the design constraints that a Wizard of Oz field study and early CHW interactions imposed, including referral advice visible first, colloquial code-mixed input, short answers readable on a phone, and hard latency, cost, and safety budgets, together with how the deployed system meets them, and the numbers behind each trade-off; and an evaluation methodology for settings without clean test sets, showing that Word Error Rate (WER) is misleading for Indic speech and describing an LLM-judge protocol aligned to, and validated against, expert preference data. Lessons learned and negative results close the paper.
Chat is not available.
Successful Page Load