Teaching LLMs to Recommend and Defer in Underrepresented Epilepsy Care
Abstract
Specialist epilepsy expertise is scarce in resource-constrained settings, making LLM-based decision support attractive for frontline clinicians managing longitudinal treatment. Such support systems must do more than apply medical knowledge: they must adapt to local prescribing practice and know when to defer. Public medical AI benchmarks are dominated by high-income clinical settings, leaving prescribing practices, medication availability, and follow-up patterns in low-resource contexts largely unrepresented. We study this problem through a multidisciplinary collaboration in Ugandan pediatric epilepsy care. The task is to predict anti-seizure medication regimens from longitudinal unstructured notes collected by local clinicians across serial visits. Standard prompting achieves non-trivial agreement with physician prescriptions, but neurologists' review of model reasoning traces shows that its errors stem from distribution-miscalibrated prescribing defaults rather than the local care environment. We introduce Manana, a non-parametric prompt-learning framework that learns how to reason about local prescribing decisions from a small patient-level training set. Manana turns observed prescription errors into an auditable prompt memory, instantiated in single-agent and multi-agent variants, and outperforms classical ML models, direct LLM prompting, and prompt-optimization baselines across two independently collected Ugandan cohorts. To make the system uncertainty-aware, we propose Bayesian prompt averaging (BPA), a Bayesian model averaging procedure over the learned prompt trajectory. This converts a sequence of learned prompts into prescription likelihoods and produces a deferral signal. On the independently collected held-out cohort, BPA improves visit-level top-3 prescription accuracy by 4-8 percentage points over the prompt-optimization baselines. More consequentially, it enables clinically meaningful selective prediction: the system can auto-handle the most confident half of cases at 95\% precision, or the most confident quarter at 99\% precision, while deferring lower-confidence cases for specialist review. These results suggest a path toward locally adapted clinical LLM systems that learn from limited site-specific data and reserve scarce specialist attention for the cases where uncertainty is highest.