PULSE: Probabilistic Uncertainty-Aware Longitudinal Simulation for EHR Trajectories
Abstract
Realistic intensive-care EHR trajectories are difficult to simulate because physiology, treatments, observation times, and patient-existence state co-evolve under sparse clinical measurement. We introduce PULSE, a probabilistic simulator family for ICU trajectories. PULSE separates treatment rollout from patient-state rollout, predicts binary patient state before continuous physiology, and updates continuous variables through gated residual dynamics around the last observed or simulated state. Probabilistic variants add event-time inputs, heteroscedastic Gaussian heads, monotone B-spline-flow residual transforms, latent severity pooling, and covariance-coupled residual groups. On a MIMIC-IV v3.1 sepsis cohort scored over a 37-variable predictive-check panel, reference PULSE reduces k=6 vector-normalized RMSE from 7.124 for a matched per-covariate transformer Monte Carlo baseline to 5.969, and improves over carry-forward from 6.633 to 5.969; paired patient-level bootstrap intervals support both differences. RMSE is a useful sanity check for aggregate scale, but not a sufficient criterion for trajectory realism: carry-forward is already difficult to beat despite generating straight-line futures. Gaussian NLL is the strongest calibrated PULSE variant by CRPS, while B-spline flows address the non-Gaussian shape of factual residuals. Within B-spline flows, event-time encodings with Δt prediction improve k=6 rollout from 8.301 to 6.635 and yield sharper, more state-dependent residual densities. In k=6 density-mixture analysis, B-spline intervals have larger across-state 90% width variation than Gaussian intervals (coefficient of variation 0.79 versus 0.48) and higher 90/95% inclusion despite slightly weaker aggregate PIT uniformity, consistent with better local skew and sharpness in some states. Similar gains over carry-forward persist on eICU sepsis validation panels, while absorbing-state timing remains a major unresolved error source. CVSim known-effect experiments show that synthetic treatment response can be learned under mechanistic ground truth, although transfer to MIMIC-seeded counterfactual tests remains weak.