Uniform Action Support Does Not Ensure Precise Policy Evaluation: A Stopped-Trajectory Sepsis Audit
Abstract
Uniform action support removes one obstacle to off-policy evaluation, but does not determine the precision of a finite-sample longitudinal estimate. We audit this distinction using fresh trajectories from a public sepsis simulator, known uniform logging over eight actions, three fixed target policies, and horizons from 1 to 16. Thirty-two independent logs contain 65,536 episodes; 196,608 separate on-policy episodes provide Monte Carlo references. Likelihood products stop at absorption rather than accumulating fictitious post-terminal factors. The prespecified primary comparison does not establish declining self-normalized importance-sampling bootstrap coverage: at the strongest policy shift and 2,048 episodes, reference-point coverage changes from 28/32 to 29/32, a difference of 0.031 with a paired 95% interval of [-0.094, 0.156]. Precision nevertheless deteriorates: median trajectory effective sample size falls from 304.6 to 10.4, median interval width grows from 0.037 to 0.391, and root mean squared error grows from 0.011 to 0.191. Ordinary and per-decision estimators do not remove the problem on this panel. The contribution is a reproducible, stopped-trajectory diagnostic of a known horizon effect, including an inconclusive coverage test and reference uncertainty. It is not a new estimator, a clinical validation, or evidence that all policy-evaluation methods fail.