FindingFrame: Auditable Longitudinal Memory from Radiology Reports
Abstract
Longitudinal oncology review requires more than extracting findings from one radiology report. It requires deciding whether two mentions on different dates describe the same clinical finding and retaining the sentence that supports each claim. We present FindingFrame, a pipeline that assigns each finding six controlled slots and a mandatory evidence span, then computes cross-report identity from normalized type, anatomy, and laterality. A large language model (LLM) reads the report; deterministic code verifies provenance, normalizes slots, and builds tracks. On 300 MIMIC-IV reports from 30 selected longitudinal patients with oncology- relevant and complex follow-up findings, FindingFrame reaches patient-averaged type, identity, and full-frame F1 of 0.859, 0.758, and 0.645. Holding the model fixed on eight development patients, the pipeline raises full-frame F1 from 0.152 to 0.771 for GPT-5 and from 0.089 to 0.599 for GPT-4o. The remaining gap is informative: 56 of 123 track splits arise when the same anatomy is recorded at different resolutions. The evaluation is also sensitive to that choice, with full-frame F1 ranging from 0.322 to 0.645 on unchanged predictions under four anatomy criteria. External chest-radiography tests preserve type F1 of 0.737 and 0.782 and evidence anchoring above 96%, but strict identity F1 remains 0.202 and 0.152. FindingFrame makes longitudinal identity reproducible after extraction and makes its failures inspectable; it does not establish clinical validity, because the reference gold is engineering-curated rather than clinician-adjudicated.