Better Than Clinician Care? A Stress Test of Off-Policy Evaluation in ICU Transfusion
Abstract
Artificial intelligence systems are increasingly proposed for treatment decisions in the intensive care unit (ICU), raising the question of how such systems should be evaluated. Treatment decisions by clinicians are assessed in randomized trials and large prospective studies. Decisions produced by learned policies are instead evaluated retrospectively, either by concordance with human experts or by off-policy evaluation (OPE). We therefore test whether \OPE{} supports a stable, credible, and clinically sensible choice among policies for red-blood-cell transfusion in the ICU. We treat clinical decision support as a reinforcement learning policy learned from offline data, and we score 21 learned, guideline-based, and planted control policies with six value estimators on two large databases under a fixed evaluation protocol. Our results show that the estimators do not agree. Every setting produces between three and six different winners, the median policy moves 8 to 12 rank positions across estimators, and a planted control takes the first rank 14 times out of 144. The strongest agreement occurs among importance-weighted estimators, but it is not independent evidence, since all four reuse the same likelihood ratios. Among the policies repeatedly ranked among the best, the database-specific median effective sample sizes are in the single digits despite thousands of held-out stays. Changing only the behaviour model changes the selected policy in 26 of 32 estimator--setting comparisons. Highly ranked policies can still recommend transfusion at normal hemoglobin levels. Retrospective evaluation therefore does not substantiate a claim that an offline policy improves on recorded clinician care.