FERMAT: An Autoregressive Clinical Foundation Model for Disease Risk and Clinical Event Prediction
Abstract
Clinical foundation models learn patterns in longitudinal health records that can be reused across prediction tasks. However, the relative strengths of representation-based and generation-based prediction across clinical tasks remain unclear. Using longitudinal electronic health records from 3.77 million patients at a single hospital, we developed FERMAT (Foundation model for Exploring Real-world Multimodal health data using Autoregressive Trajectory modeling), a 65-million-parameter clinical foundation model. We compared predictions from task-specific predictors trained on patient representations with those derived from generated clinical event sequences. Representation-based prediction performed better in the disease-risk comparisons. In detailed analyses of five-year chronic kidney disease, fatty liver, and diabetes risk in a test cohort of 10,000 patients, representation-based predictors exceeded generation-based estimates by 0.098-0.115 in observation-weighted AUROC and had lower probability error. Generation-based prediction, in turn, performed better in predicting diagnoses, prescriptions, and procedures recorded on the same day, a task motivated by the frequent presence of multiple events with the same date in hospital EHRs. Using 50 generated trajectories per patient, it achieved a Recall@10 of 71.0\%, compared with 55.7\% for representation-based prediction, with a paired improvement of 15.31 percentage points (95\% CI, 13.92-16.79; n = 1,862). These findings demonstrate task-dependent strengths of representation-based and generation-based applications within a single clinical foundation model, providing an empirical basis for exploring how the two approaches can be used individually or combined to address clinical prediction needs.