Cross-System Reuse of an EHR Foundation Model for Prediabetes Trajectory Phenotyping
Abstract
Patient health records accumulate over years into long and highly individual histories that are difficult to summarise and model. Despite this complexity, characterising shared trajectories within these histories could improve understanding of disease heterogeneity and progression. Electronic Health Record (EHR) foundation models solve this challenge by encoding longitudinal clinical histories into compact, lower-dimensional patient representations, or embeddings, that can be reused across analytical tasks. However, their ability to transfer across health systems and populations remains an open question. In this work, we used heterogeneity in progression from prediabetes to type 2 diabetes as a clinically meaningful use case of whether a US-trained EHR foundation model could transfer zero-shot to a UK health system. We applied the Mamba-based EHR foundation model released by Stanford Shah Lab to full longitudinal UK trajectories from adults with prediabetes. This created a challenging transfer setting across national health systems, care settings, and clinical coding vocabularies. Patient-level embeddings generated from these records were then clustered to identify distinct clinical trajectory phenotypes. Despite this substantial shift, the embeddings recovered six clinically interpretable patient groups with significantly different rates of progression to type 2 diabetes, and distinct patterns of cardiometabolic burden, mortality, glycaemic change, and healthcare use. We additionally evaluated embedding quality by predicting future prediabetes from EHR data available 1, 4, 10, or 15 years before the prediabetes-defining index date. Compared with a logistic-regression classifier using bag-of-codes features from the same histories, the embedding-based classifier performed slightly worse at 1 year, similarly at 4 and 10 years, and slightly better at 15 years, suggesting that the embeddings retained metabolic signal in more distant clinical history. These results suggest that, despite the challenges of zero-shot transfer across healthcare systems, pretrained EHR foundation models can still preserve clinically meaningful information that is useful for exploratory patient sub-phenotyping. This could be particularly valuable when local training data or resources are limited. With appropriate vocabulary harmonisation, pretrained models may be applied to longitudinal EHRs with less disease-specific feature engineering, potentially streamlining the study of patient trajectories for disease understanding, patient care, and therapeutic discovery.