Matched Budgets or Meaningless Numbers: Baseline Under-Specification as the Weak Link in Patient World Model Validation
Abstract
Patient world models—simulators, digital twins, and synthetic control arms trained to predict individual clinical trajectories—are validated almost exclusively on outcome fidelity: how closely a model's predictions match a held-out patient's real trajectory. Left unspecified, almost without exception, is the comparator's information budget: what data, and under what protocol, the baseline was allowed to see. Because a reported advantage over that baseline is a function of this unstated methodological choice, current validation practice cannot distinguish a genuinely better world model from a more favorably configured comparator. We formalize this gap as six binary pre-specification items—real historical comparator vs. held-out split; comparator information budget stated and matched; uncertainty or calibration reported; an abstention or selective-prediction rule; positivity/overlap diagnostics; and public code with cohort definitions—and use them to audit 24 recent papers on patient world models, virtual and synthetic control arms, counterfactual outcome prediction, and EHR foundation models. We publish our search protocol, query strings, and inclusion criteria, and report a per-paper, six-column scorecard rather than a single aggregate score, so any reader can re-score the sample. As a solo-rater audit, we substitute full methodological transparency for claimed inter-rater agreement. The paper's primary artifact is a numbered pre-specification checklist that trial and simulator teams can adopt directly, before results are ever compared.