Factual Validation Can Select the Wrong Intervention Kernel: A Sepsis-Simulator Audit
Abstract
Which validation score should select a model that will be queried under a different treatment policy? In a public sepsis simulator, we first show that better factual prediction can accompany worse intervention prediction: under strong treatment selection, pooling improves factual Brier loss over standardization by 0.04641 but worsens randomized-target kernel error by 0.01878. We then test a remedy without providing intervention outcomes to the selector. On 64 fresh states, 12 data replicates and three assignment regimes, ordinary and importance-weighted factual validation select from the same six fitted candidates. Selections are frozen before independent scoring references are generated. Weighting reduces selected-kernel error from 0.03221 to 0.01344 at assignment strength 0.8, and from 0.06069 to 0.04162 at 0.95; adjusted paired intervals exclude zero. The selectors coincide under randomized assignment. These two studies comprise 1,507,328 simulator transitions. The contribution is a reproducible evaluation and selection audit using established estimators, with known-propensity and fixed-support limitations, not a new causal identification result or evidence of clinical benefit.