Sharp Characterization of Bias in Post-Bandit Inference
Lisu Wang ⋅ Yilun Chen ⋅ Jiaqi Lu
Abstract
Bandit algorithms generate data for downstream inference, but adaptive sampling biases post-bandit sample means. For stable index algorithms including UCB1, we derive sharp characterization for both the raw sample-mean bias and the expected $Z$-statistic. The bias is governed by an index-dependent effective exploration rate: under UCB1, standardized bias for any arm that is not uniquely optimal decays only as $1/\sqrt{\log T}$. The result reveals the algorithmic origin of bias in post-bandit inference, and highlights when hypothesis testing is most severely affected.
Chat is not available.
Successful Page Load