TRIAGE-WM: Protocol-Balanced Reliability Audits for Selective Clinical Trial Simulation
Abstract
Patient world models are increasingly discussed as engines for virtual trial arms and counterfactual treatment simulation, but a clinically useful simulator must know when a requested rollout is weakly supported. We introduce TRIAGE-WM, a model-agnostic audit that records three signals along an intervention path: ensemble disagreement, treatment-overlap surprise, and state-action support distance. We study two ways to scalarize this vector—an amplification-aware percentile score and a residual-calibrated risk score—and introduce protocol-balanced selection, which requires the same retained coverage within each intervention arm. Across five seeds on a nonlinear longitudinal structural benchmark, the amplification-aware score reduces endpoint MAE at 60% global coverage by 56.2% in-distribution and 47.7% under covariate shift. The same score, however, fails to rank error in a PK-PD tumor-growth stress test and becomes anti-informative under a hidden treatment-effect shift. A learned residual score reverses this pattern: it works well on the tumor simulator but not reliably on the generic benchmark. More importantly, global abstention creates severe protocol imbalance: at nominal 60% coverage, the retained-fraction gap between trial arms reaches 87.0 percentage points on the generic benchmark and 83.9 points on the tumor benchmark. Protocol balancing removes this gap but can erase much of the apparent error gain. The results argue that selective clinical simulation should be evaluated as a trial-level policy, not merely by a single confidence score: reports should expose component diagnostics, arm-stratified coverage, and mechanism-shift negative controls. All experiments are synthetic; no patient data are used.