CARE: A Conformal Safety Layer for Medical Summarization
Bridget Lin ⋅ Suhana Bedi ⋅ Anson Zhou ⋅ Chloe O Stanwyck ⋅ Jenelle A Jindal ⋅ Sanmi Koyejo ⋅ David Stutz ⋅ Nigam Shah
Abstract
Large language models (LLMs) are increasingly used for medical summarization but can omit medically important information and introduce unsupported claims. Existing error-detection methods produce heuristic or uncalibrated scores, providing no formal control over missed errors and no principled way to trade off safety against clinician review burden. We introduce Conformal Assessment for Risk Evaluation (CARE), a post-hoc risk-control layer that adds calibrated omission and hallucination annotations without retraining the summarizer. CARE uses scalar conformal risk control (CRC) to bound the probability that a document contains an unflagged hallucinated sentence, and Learn-Then-Test fixed-sequence testing (LTT-FST) to control the expected fraction of true omissions left unsurfaced for review. Omission surfacing depends jointly on sentence importance and summary coverage, creating a two-dimensional threshold-selection problem. We show that calibrating one gate and then imposing an uncalibrated second gate can invalidate risk control. At $\alpha=0.15$, LTT-FST achieves the lowest review workload among the risk-controlled omission baselines. Across five medical summarization tasks, mean held-out violation over 100 random calibration/test resplits remains at or below the target for both controllers. Review workload varies by task and is highest for long, high-compression documents. In a preliminary study of 75 clinician reviews, omission coverage increased by 28.6 percentage points. Overall, CARE provides a principled framework for targeted, risk-controlled review of LLM-generated medical summaries.
Chat is not available.
Successful Page Load