Variational Generative Thompson Sampling for Nonstationary Contextual Bandits
Abstract
Nonstationary contextual bandits require learning not only the reward relationship between contexts and actions, but also how this relationship evolves over time. We propose Variational Generative Thompson Sampling (VG-TS), which learns a latent representation of the evolving reward process where flexible reward modeling is combined with simple temporal dynamics. VG-TS uses a variational encoder, a Gaussian process prior over latent states, and Thompson sampling over the posterior of the current latent state to generate coherent samples of the current reward environment. Experiments on a nonlinear synthetic benchmark and the MIND news recommendation dataset demonstrate that learning a representation in which nonstationarity is simple can provide an effective approach for uncertainty-driven decision making.