Bandit Simulation for Average Reward Inference
Abstract
Multi-arm bandit algorithms are increasingly used in online platforms, clinical trials, and social science experiments, but valid statistical inference on their performance remains an open challenge. After deploying bandits, a natural question is whether one can construct a confidence interval for its mean reward and assess whether it reliably outperforms a baseline policy. More broadly, one may wish to assess a new candidate algorithm's expected reward before committing to deployment. Standard inference methods break down because bandit algorithms introduce complex dependencies in the collected data, leading to inflated Type-I error. Moreover, existing inference methods for bandit data only apply to estimands such as the mean reward under a fixed action, which do not depend on the data-collection algorithm. We propose Bandit Simulation for Inference (BSI), a framework that fits a simulator of the bandit environment from on- or off-policy data and uses it to estimate the mean reward under any evaluation policy, including adaptive blackbox algorithms. BSI formally propagates simulator estimation error into the confidence interval construction, requires only weak exploration assumptions on the behavior policy, and avoids importance weighting. We prove that BSI yields asymptotically valid confidence intervals, and demonstrate empirically that it maintains nominal coverage in settings where standard off-policy evaluation methods fail.