Should Register Foundation Models Be Part of Pandemic Preparedness?
Abstract
Foundation models trained on population-scale event sequences are designed to capture general patterns in individual trajectories for prediction, simulation, and other downstream analyses. Whether this generality would help at the start of a new crisis remains unknown. To test preparedness, the model must be evaluated on a crisis that was absent from its pretraining data. To isolate what the model could carry into a new crisis, we trained four register foundation models on nationwide Swedish register data truncated at successive cutoffs before Sweden's first confirmed COVID-19 case, and five, ten, and seventeen weeks later. The models shared the same architecture, hyperparameters, and training budget. At each cutoff, we used generative rollouts to estimate 30-day risks of death, hospitalization, and sickness absence among laboratory-confirmed cases in the following 60 days. We also scored every model on the cases of each later cutoff. We evaluated the foundation models against a pre-pandemic age- and sex-specific incidence baseline available at every cutoff, and against supervised models once sufficient labels had matured. With no pandemic events in pretraining, the model reached AUROCs of 0.85 for death and 0.82 for sickness absence, compared with 0.88 and 0.77 for the incidence baseline. After five weeks, before enough 30-day outcomes had matured to fit a supervised model, the foundation model's death AUROC exceeded the baseline. Its death Brier score was lower than the baseline at both early cutoffs. Brier scores worsened sharply for hospitalization and sickness absence at the five-week cutoff, and for death at the ten-week cutoff. By seventeen weeks, the model had the highest AUROC point estimate for sickness absence but was surpassed by supervised models for hospitalization. Scored on the same later cases, models exposed to pandemic data did not consistently improve discrimination over the pre-pandemic model and often had substantially higher Brier scores. Additional early-pandemic pretraining therefore did not reliably improve predictive performance.