Temporal Hallucinations in Medical Report Generation
Gefen Dawidowicz ⋅ Malak Fares ⋅ Ayellet Tal
Abstract
Current evaluation metrics for automated medical report generation fail to assess temporal reasoning, obscuring critical errors in disease progression. To address this, we introduce a dual-stage LLM-as-a-judge framework designed to explicitly quantify temporal hallucinations. Applying this tool to state-of-the-art models on MIMIC-CXR reveals severe deficiencies: single-image models hallucinate temporal comparisons in up to 89% of cases, and even prior-aware longitudinal models fail in approximately 50% of cases, predominantly via Temporal Omissions. Ultimately, this work exposes a critical performance gap and provides the essential benchmarking tool needed to develop genuinely accurate longitudinal models.
Chat is not available.
Successful Page Load