Adaptive Top-$n$ Retained-Set Exploitation in Weakly Separated Bandit Regimes
Advaith Anand ⋅ Abhishek Umrawal
Abstract
Many bandit methods exploit by repeatedly choosing the action that currently appears best. This strategy can be brittle when leading actions are weakly separated, as small estimation errors or short-term fluctuations can cause premature commitment while several alternatives remain competitive. We study Adaptive Top-$n$ Retained-Set Exploitation, which retains top-ranked actions, chooses the retained-set size online using the current estimator state, and distributes exploitation probability within that set. We evaluate fixed and adaptive variants across Gaussian bandit regimes with varying separation and nonstationarity, along with different within-set weighting rules. Adaptive retained-set exploitation produces consistent gains over fixed retained-set selectors in favorable weakly separated regimes, with the magnitude of those gains also depending on the weighting rule. These findings show that robustness to near-tied scores can come not only from estimation but also from how scores are converted into actions.
Chat is not available.
Successful Page Load