SurvBench: A Standardised Benchmark for Multi-Modal Electronic Health Record Survival Analysis
Abstract
In recent years, deep learning for time-to-event or survival modelling has allowed clinical AI models to predict risk dynamically across different care settings. Although originally proposed for epidemiological or survey data, more recent methods have been applied in highly sampled shorter horizons in the hospital, including critical care. However, comparing these methods is difficult due to somewhat subjective decisions on dataset preprocessing, censoring assumptions, survival labelling, and binning. Here, we present SurvBench, an open-source benchmark that standardises large EHR data for survival analysis. We include four emergency/critical-care databases (MIMIC-IV, eICU, MC-MED, HiRID) and four modalities, time-series vitals and laboratory values, static demographics, International Classification of Diseases (ICD) codes, and radiology reports. The preprocessing pipeline is configurable, provides explicit missingness masks, and supports landmarked prediction. The benchmark includes both single-risk and competing-risks outcomes, and we provide alternative choices for censoring. The benchmark also includes five baseline models, Cox proportional hazards, DeepHit, Dynamic-DeepHit, DySurv, and a survival transformer. Across datasets, temporal information was consistently important: removing time-series features reduced time-dependent concordance by 0.104–0.243, whereas removing static features changed it by only 0.003–0.008. Multimodal gains were dataset-dependent, improving landmarked competing-risk prediction on MC-MED but not mortality prediction on MIMIC-IV. Cross-dataset eICU-to-MIMIC-IV experiments further showed that model rankings changed with the amount of available target-domain data, highlighting the importance of evaluating survival models across preprocessing choices, prediction settings, modalities, and domains.