Longitudinal Clinical Notes for Hospital Course Summarization: A Benchmark of Input Construction and Claim-Level Evaluation
Abstract
Hospital-course summarization requires models to synthesize long, repetitive, and potentially incomplete clinical records, yet prior benchmarks often use discharge summary content or substantially reduced clinical notes as input source document to generate hospital course summary. We examined whether fine-tuning gains observed under these formulations persisted with longitudinal source documentation. We constructed four matched input variants for 860 MIMIC-III admissions: Full-Context Notes, Section-Filtered Notes, 24h Pre-Discharge Notes, and discharge-summary content excluding the target Brief Hospital Course. We char acterized each variant by length, lexical redundancy, and support for clinically salient reference claims. Fine-tuned Llama 1B and 8B models were compared with their untuned counterparts and larger zero-shot open-weight models using ROUGE, BERTScore, and a validated claim-recovery metric. Full-Context Notes provided the greatest reference support, while section filtering retained most supported in formation with shorter inputs. Fine-tuning gains varied by input construction and model capacity: Llama 8B improved on full-context and section-filtered inputs, but gains were smaller or absent when inputs omitted more reference information. Larger zero-shot models benefited more consistently from full context. These results show that input construction changes both the information available to models and conclusions about fine-tuning, model capacity, and summary quality.