Simulating Patient Futures: A 2.7-Million-Patient World Model Across Health Networks
Abstract
Here is the raw text of the abstract with all citation references removed:
Generative models of patient timelines can answer questions care teams actually ask—who develops hypertension next year, who is readmitted, how many visits to staff—by simulating the next year and counting outcomes. Almost every published model of this kind is trained on one health system (often ICU) and scored on next-token accuracy.
We train an ETHOS-family decoder from scratch on 2.7M patients from several U.S. outpatient networks (not one hospital). We then test it in two places. Our representation baseline is CLMBR-T, a published single-system EHR foundation model; we use its public pretrained weights as released, with no fine-tuning of CLMBR-T itself.
Held-out 100k. Same networks, never used in training. A simple linear probe on our frozen model beats that frozen CLMBR-T on all 18 disease and visit tasks we can compare fairly (about +0.12 AUROC on average). The same model, with no fine-tuning, can also simulate 20 possible futures per patient and score the year ahead: 0.89 AUROC for new hypertension and 0.91 for new high cholesterol, versus 0.51 for MedGemma-4b given the same records and told which codes to look for.
EHRSHOT. Stanford’s benchmark—a different health system the model never saw. We remap the records into our vocabulary and again use only a frozen probe (no retraining). On the eight outpatient tasks we are about tied with the same frozen CLMBR-T weights (mean AUROC 0.72 vs. 0.73), despite missing most of their diagnosis codes.
Two limits. A gradient-boosted model using only age and past visit counts is close on new diagnoses and beats us on readmission. And on these tasks, the model’s learned representation carries most of the value; full simulation helps mainly where a single next-token guess fails.
Takeaway. Training on many real-world outpatient networks is the cheap lever. A linear probe on a broad-corpus model is a drop-in upgrade over a site-native one; reserve simulation for cases where sampling is the only way to recover the signal.
Authors: Tapan Shah, Janak Ramachandran