Grounding Language Models with Symbolic Knowledge and Selective Prediction for Safer Clinical Trial Matching
Abstract
Motivation. Clinical trial matching is a high-impact but cognitively demanding task: it requires interpreting heterogeneous, unstructured clinical evidence against dense, multi-criterion eligibility protocols, and errors here directly gate a patient's access to potentially life-extending treatment. While large language models (LLMs) have shown promise in automating this process, their reliability remains insufficient for high-stakes decision support — unconstrained probabilistic reasoning, hallucinated entities, and inconsistent handling of uncertainty are recognized failure modes, and real-world evaluations highlight a substantial gap between retrospective performance estimates and prospective deployment: systems reporting near-expert sensitivity in retrospective meta-analyses (~90%) achieved only ~32% sensitivity when publicly deployed matching tools were tested prospectively on real patients (Gueguen et al., 2025). We ask whether combining symbolic grounding with an explicit selective-prediction mechanism can reduce this generalization gap while remaining safe to deploy alongside human experts. Method. We propose a triage-aware neuro-symbolic framework that separates perception from reasoning: unstructured clinical text is processed by an LLM (Llama 3.1-70B) to extract structured variables, which are grounded to canonical concepts via a symbolic knowledge graph linked to open public vocabularies (NCIt, HGNC, OncoTree) and enforcing variant-level identity constraints — e.g., a KRAS G12C trial cannot match a KRAS G12V patient. Eligibility is then evaluated with deterministic, per-criterion rules encoded directly from full trial protocols, producing auditable explanations rather than an opaque score. A deterministic triage layer then performs selective prediction: a confidence score, computed from evidentiary completeness (missing inclusion/exclusion evidence) and detected conflicts between hard criterion failures and missing evidence, is thresholded into three tiers — automated screening, review, or mandatory human evaluation — so uncertainty is handled transparently rather than silently propagated into a final decision. Results. Across 77 lung-cancer patients and 17 concurrent thoracic trials (1,309 patient–trial evaluations), the system achieved a false-positive rate of 0.0% and ranked the correct trial first for 55 of 55 ground-truth- eligible patients (100%, 95% CI 93.5–100%). The triage layer judged 49% of cases suitable for full automation (49% coverage); the automated subset had a 0% false-positive rate, with the remaining 51% deferred to clinician review. In an ablation subset (n=35), the full system was compared against an LLM- only pipeline that re-reasons over each trial separately: Top-1 accuracy rose from 0.80 to 1.00, Recall@3 from 0.87 to 1.00, false-positive rate fell from 0.15 to 0.0, and processing time from ~12 min to <1 min per patient — consistent with a single extraction pass reused across all 17 trials rather than per-trial LLM re- reasoning. The framework transferred without modification to an independent genitourinary cohort (30 patients, 1 trial), where it introduced no additional ranking or false-positive errors; across the combined 107- patient cohort spanning both tumor types, Top-1 accuracy and Recall@3 remained 100% with a 0.0% false- positive rate, providing evidence of transfer across tumor types without domain-specific modification. Significance. Grounding LLM outputs in symbolic, auditable reasoning, combined with principled abstention, offers a pathway toward safer real-world deployment, directly addressing the prospective- validation gap that has limited prior work. These results suggest that reliable high-stakes LLM systems may benefit less from forcing confident end-to-end predictions than from explicitly separating perception, symbolic reasoning, and abstention — a principle applicable to any domain where LLMs act under partial evidence and errors are asymmetric in cost. Ongoing work adds formal temporal reasoning (Allen interval algebra), patient-specific benefit-aware trial ranking via survival modeling, and prospective deployment alongside routine clinical care.