Quantifying the Silent Failure Modes of Generative Electronic Health Record Models
Abstract
Generative foundation models for electronic health records (EHRs) have shown promising zero-shot abilities to predict medical outcomes by autoregressively sampling future patient trajectories. However, error rates in generating these timelines and the methods used to handle them are rarely reported, potentially masking critical reliability risks for clinical deployment. This paper provides the first quantitative evaluation of patient trajectory error rates, their underlying failure modes, and their downstream impact on performance metrics. Using the MIMIC-IV and MIMIC-IV-ED datasets, we compare two Gemma-3-270M-based models trained on distinct modalities: a structured EHR model trained from scratch with a custom vocabulary, and a natural language model fine-tuned from a general-domain checkpoint. Our evaluation across five common medical outcome prediction tasks reveals significant, task- and modality-dependent error rates ranging from 19.25\%--62.08\%. We identify a clear divergence in failure modes: the structured model has high syntactic error rates, whereas the natural language model is syntactically more robust but is bottlenecked by output context-length limitations. Furthermore, we demonstrate that the choice of error-handling strategy, such as dropping invalid trajectories or assigning priors, significantly alters standard performance metrics such as AUROC by up to 0.262 and can even silently exclude high-risk patient cohorts from evaluation. These findings expose a vulnerability in autoregressive EHR prediction models, underscoring the urgent need for transparent error reporting and motivating future work on efficient constrained decoding, longer-context models, or instruction-tuned question-answering frameworks.