RETINA-SAFE and ECRT: Evidence-Conditioned Hallucination Risk Triage for Medical LLMs
Zhe Yu ⋅ Wenpeng Xing ⋅ Meng Han
Abstract
Medical large language model (LLM) safety depends on how an answer relates to the available evidence, not on the input scenario alone. A model can safely reject a false premise or defer when evidence is incomplete; the same scenario becomes unsafe when the model asserts a contradicted or unsupported conclusion. We introduce RETINA-SAFE, a 12,522-item diabetic retinopathy benchmark that explicitly separates three evidence relations from model-specific response-safety labels. We also introduce ECRT, a two-stage detector that holds the generated response fixed and contrasts the backbone's logits and hidden states with retinal evidence present (CTX) and removed (NOCTX). Stage 1 detects unsafe responses, and Stage 2 attributes contradiction versus evidence-gap risk. Across five patient-disjoint grouped holdouts, ECRT reaches $0.8142 \pm 0.0045$ balanced accuracy and $0.8845 \pm 0.0048$ AUROC on Llama-3-8B, exceeding FactoScope by 0.0632 in balanced accuracy (95% CI: [0.0512, 0.0754]). On Qwen2.5-7B, ECRT reaches $0.8435 \pm 0.0041$ balanced accuracy and $0.9120 \pm 0.0042$ AUROC, a 0.0615 gain over FactoScope. At a validation-selected operating point targeting at least 95% unsafe-response recall, the estimated review rate falls from 82.6% with perplexity to 62.8% with ECRT on Llama-3-8B. Text shortcuts, negative controls, counterfactual interventions, and a source-held-out evaluation show that the gain tracks the evidence–response relation within the tested retinal setting.
Chat is not available.
Successful Page Load