An Empirically Sharper Estimator for Confounding-Robust Reinforcement Learning
Abstract
Offline policy evaluation is vulnerable to unobserved confounding, motivating sensitivity analysis that returns a range of policy values under a bounded-confounding model. In one-step problems, sharp bounds are characterized by conditional quantiles, and recent work showed that KCMC can improve the finite-sample lower bound relative to a quantile-based benchmark using the same basis-function span. We extend this comparison to finite-horizon reinforcement learning by introducing KCMC fitted-Q evaluation and iteration. We prove a two-sided result: for matched robust backups, KCMC gives a lower bound no smaller and an upper bound no larger than the quantile-based estimator, and we identify conditions under which this advantage survives fitted-Q recursion. Controlled experiments show that the improvement is often strict and can remain substantial after recursive policy evaluation. Thus, in the matched setting, KCMC is theoretically no looser and can be empirically sharper for confounding-robust reinforcement learning.