Auditing ICU World Models: Support, Predictive Utility, and Intervention Response
Abstract
Action conditioning alone does not make a clinical sequence model a credible simulator of alternative care. We present a staged rollout credibility audit that asks, for each proposed query, whether comparable decisions provide empirical support, recursive rollouts add factual utility beyond persistence and observation anchors, and predictions respond coherently to controlled changes in acquisition history or intervention inputs. We apply it to an action-conditioned latent model and two external model recipes, including a paired loss ablation, using a cohort of \NUM{73{,}073} MIMIC-IV ICU stays. Support is sharply query-dependent. In the latent model, much long-horizon skill is recovered without the rolled latent, and intervention deletions passing the assignment-score screen produce small responses, with several stable contrasts opposing prespecified directional hypotheses. Factual skill and response behaviour diverge across recipes. A known-truth longitudinal simulation shows that observed support and predictive skill can coexist with large effect error under hidden confounding. The audit maps query-, horizon- and model-specific failure boundaries, but does not identify treatment effects from the clinical records. The evaluated system is therefore a world-model candidate rather than a validated simulator for action selection.