Performance Incentivization for Multi-armed Bandits with Stochastically Strategic Agents
Edward Han ⋅ Alexander Wang ⋅ Fatemeh Ghaffari
Abstract
We study bandits in which a selected strategic agent chooses costly effort that can raise or lower the mean reward, while the platform observes only a noisy Bernoulli outcome. The platform must encourage high performance without losing the usual protection offered by honest agents. For fixed stationary means, we show that symmetric UCB adapts to the best delivered mean, favors higher delivery, and treats equally performing top agents symmetrically. This yields an honest-agent reward guarantee and, when capacities are public and at least two agents share the highest capacity, a high-performance equilibrium. Private capacities are harder: bids must be estimated from noisy samples, and an agent can bid high before reducing its effort. Our monitored empirical second-price mechanism combines statistical bidding, continuous monitoring, and an EXP3 fallback. On every fixed scalar instance, it gives an additive equilibrium against arbitrary feasible nonanticipating unilateral deviations, earns the second-highest capacity benchmark up to $O(n^{2/3}(\log n)^{1/3})$, and preserves the honest benchmark under arbitrary behavioral play. For contextual problems, independent \UCB{} instances give exact guarantees on finite profile spaces. For i.i.d. continuous profiles in dimension $D:=Kd$, static grids give robustness error $\widetilde O(n^{(D+1)/(D+2)})\)$; an explore-then-commit variant also gives contextual incentive properties and, under explicit top-stability assumptions, a restricted public equilibrium. Simulations illustrate the incentive response and the finite-horizon cost of robustness. Private contextual incentives remain open.
Chat is not available.
Successful Page Load