Semantically Identical, Clinically Opposite: A Negative Result on Semantic Caching in a Deployed Maternal-Health Assistant
Abstract
We run a human-in-the-loop AI assistant for Auxiliary Nurse Midwives and other frontline health workers managing high-risk pregnancies across eight Indian states, supporting four Indic languages and their code-mixed variants. Cost and latency limit its scale, and unreliable connectivity limits where it can be used. Both point to the same fix: reuse verified answers to equivalent questions, with an on-device cache that escalates only novel queries. We spent six months evaluating whether this is safe, entirely offline, replaying production history through the same lookup path. It is not. Semantic caching assumes similarity implies answer interchangeability. In clinical guidance it does not, because the features that determine the answer are the ones sentence-level similarity discounts. Across six embedding models, altered clinical numbers outranked genuine paraphrases for most benchmark anchors. Safe thresholds collapsed the hit rate. Trained rerankers scored 88% in cross-validation but 44% on unseen clinical topics. LLM canonicalization into hashed keys matched half the queries but produced clinical violations in 44.8% of matches. We report the negative result and the failure mechanism rather than the system we intended to build.