Prospective Pilot Deployment of a Large Language Model System for Radiologically Relevant Clinical History Summarization
Abstract
Clinical context is essential for accurate radiologic interpretation and appropriate imaging protocol selection. However, clinical histories accompanying imaging orders are often incomplete or nonspecific (e.g., "pain", "cancer eval", etc.). We developed a large language model system that extracts and summarizes radiologically relevant clinical information from the electronic health record, and prospectively deployed it into a routine academic radiology workflow using CDS Hooks triggers, FHIR-based note retrieval, and PACS-integrated display of the generated history alongside the imaging study. During the pilot, the system generated 58,063 clinical histories. Radiologists accessed summaries for 1,613 examinations and submitted 383 structured ratings on a self-selected subset of these. Of these ratings, 87.47% (335/383) indicated improved clinical context relative to the referring clinician's indication and 87.73% (336/383) reported no clinically incorrect, missing, or misleading content. However, 3.39% (13/383) flagged errors with potential negative effect on interpretation or report impression, and free-text feedback surfaced hallucinations and formatting failures absent from prior retrospective evaluation. We report the design and initial pilot results of this deployment and discuss observed limitations and areas for future improvement.