Heteroscedastic Gaussian Process Q-learning for Sample-Efficient Experimentation
Abstract
In process experimentation, chemical reactors, and bioprocesses, reinforcement learning agents may only have access to a few hundred trajectories. Each is a real experiment and therefore, the employed algorithm must use them efficiently. Furthermore, any tuning of algorithm hyperparameters can further waste experimental budget. Standard neural Q-learning offers no closed-form uncertainty estimates, and requires additional tuning to achieve satisfactory performance. Non-parametric alternatives can improve sample-efficiency through improved function approximation, however existing methods either fix a constant step or target exploration regret. We propose Heteroscedastic Gaussian Process Q-learning (HGPQ): a closed-form Bayesian gain, obtained by projecting both moments of the value posterior under the distributional Bellman operator, replaces the scalar learning rate of prior Gaussian Process Q-learning, with a heteroscedastic variance update making the gain itself state-action-dependent. Under a fixed, small interaction budget across chemical process-control and biological-growth case studies, we compare HGPQ against specialized low-data RL algorithms and standard vanilla RL algorithms. HGPQ matches or improves on every low-data comparator, and deep actor-critic algorithm.