Learning to Harness Data Center Flexibility with Adaptive Bandits for Sustainable AI
Abstract
The rapid growth of large-scale AI workloads in data centers has placed increas- ing pressure on power grids in recent years. Since power systems must con- tinuously balance supply and demand, there is growing interests in leveraging data-center workload flexibility as a grid service. We propose a contextual rest- less multi-armed bandit (CRMAB) framework in which a grid operator requests load reductions without observing internal job-scheduling decisions. Under index- ability guarantee, each data center or physical machine is modeled as a Markov decision process (MDP) over a cyclic virtual-machine (VM) job queue, with un- known rewards and transition dynamics learned online using Thompson sampling and Whittle-index policies. To improve learning under sparse and noisy observa- tions, the framework augments an adaptive Thompson–Whittle (TW) policy with domain-informed transition priors and gated prior mixing. In baseline experi- ments, the best adaptive refined variant achieves 91.4% of the oracle reward after 100 rounds and 96.8% after 1,000 rounds. Across a 16-setting stress test spanning different state-space sizes and levels of contextual noise, the best refined vari- ant consistently outperforms the original TW policy with high confidence while remaining competitive with EXP4. A graph-based prior further incorporates data- center hardware constraints, including computing-resource limits. Overall, the re- sults demonstrate the economic potential of data-center flexibility as a grid service and highlight the importance of high-quality, open-source AI workload datasets for developing and evaluating such services