Should Register Foundation Models Be Part of Pandemic Preparedness?
Abstract
Foundation models trained on population-scale event sequences are designed to capture general patterns in individual trajectories for prediction, simulation, and other downstream analyses. Whether this generality would help at the start of a new crisis remains unknown. Evaluating beyond the training window does not answer that question if the crisis is already present in pretraining. We tested this by training three register foundation models on nationwide Swedish register data truncated at successive cutoffs: before Sweden’s first confirmed COVID-19 case, five weeks later, and seventeen weeks later. The models shared the same architecture, hyperparameters, and training budget. At each cutoff, we used generative rollouts to estimate 30-day risks of death, hospitalization, and sickness absence among laboratory-confirmed cases in the following 60 days. We evaluated the foundation models against a pre-pandemic age- and sex-specific incidence baseline available at every cutoff, and against supervised models once sufficient labels had matured. With no pandemic events in pretraining, the model reached AUROCs of 0.85 for death and 0.82 for sickness absence, compared with 0.88 and 0.77 for the incidence baseline. After five weeks, before enough 30-day outcomes had matured to fit a supervised model, the foundation model’s death AUROC exceeded the baseline. Its death Brier score was lower than the baseline at both early cutoffs. By seventeen weeks, it had the highest AUROC point estimate for sickness absence but was surpassed by supervised models for hospitalization. At that cutoff, rollout Brier scores exceeded 0.4 for both outcomes. The model therefore provided useful early risk ranking, but gains were outcome-specific, and the probability estimates at the later cutoffs were not reliable enough for direct use.