Clinical Retrieval at Scale: Decomposed RAG in a Real-World Multi-Agent Telehealth Deployment
Tony Y Sun ⋅ Bennett Mountain ⋅ Gunnar Bell ⋅ Jaisal Friedman ⋅ Anqi Lu ⋅ Manan Shah ⋅ Michael Shick ⋅ Cyn Liu ⋅ Rumana Rashid ⋅ RuiJun Chen ⋅ Hashem E Zikry ⋅ Cian O Hughes
Abstract
We operate an AI-enabled asynchronous telehealth service in which retrieving relevant clinical context from a patient's longitudinal record is integral to every response. Patients message licensed clinicians about acute and chronic concerns over multi-turn threads, and a multi-agent orchestration framework serves those threads, powering 40k patient threads across 20k unique patients per month. Our patient-facing orchestrators autonomously conduct intake and history-taking, while our clinician-facing orchestrator enables licensed in-house clinicians to care for 10$\times$ the patients per day that they could unassisted, without compromising quality. Though methods for FHIR retrieval have been described, few describe it within a production-facing system delivering care to patients at our scale. Between January and August 2026, more than half of all clinician messages carried AI-drafted content that a licensed clinician reviewed and edited before sending. The service's first-generation retrieval tool is a monolithic design that bulk-fetches a question-agnostic snapshot of the patient record under a fixed token budget and asks a language model to filter it. We rebuilt that layer as decomposed retrieval, replacing the single tool with one query-targeted tool per clinical concept: allergies, medications, conditions, vital signs, laboratory results, procedures, encounters, and immunizations. Each tool runs hybrid lexical and dense search with reciprocal rank fusion over an index of FHIR R4 resources. A cached ``default packet'' of safety-critical data is injected into every turn. We describe the architecture, tool schemas, and the authorization and consent filters compiled into every query. We then describe the human-in-the-loop controls and report offline retrieval and answer-quality results for both generations on a benchmark of real-world clinical questions over public de-identified records.
Chat is not available.
Successful Page Load