Replica Exchange Reinforcement Learning
Jeremy Curuksu
Abstract
Reinforcement Learning is limited by the sampling inefficiencies of its inherent "explore-exploit" search algorithm. Enhanced sampling techniques are often used to search multi minima state-action and reward spaces, such as proximal/group- relative policy optimization, but time-linear policy regularization offers limited options to address the complex tradeoffs between exploration and exploitation. In this paper, we formalize parallel tempering methods known to improve sampling of complex systems in statistical physics, and extend this formalism to control theory. The learning process, which we call Replica Exchange Reinforcement Learning (RERL), maintains an ensemble of policies (replicas) each optimized under a different parameterized family of reward objectives, and allows periodic swaps of policy functions between replicas. Explicit exchanges promote the exploration of new policies compatible with a given reward objective without loosing the related policies which led to the exchange, continuing to sample and evolve locally related policies in parallel. We illustrate the improved sampling efficiency of RERL on several environments and reward objectives: on control problems with parallel tempering of random noise and only two replicas, RERL increases win rate by 25% with probability of following the optimal policy $p^* = 0.98$ vs. $p^* = 0.80$ without RERL; on verifiable math and coding tasks, RERL tempering of KL deviation from reference language models with only two replicas increases win rate across all datasets (GSM8K, MATH, MBPP) from 14% to 24% (+67%) for Qwen2.5-1.5B-Instruct and from 61% to 67% (+9%) for GPT-4o-mini. Our results demonstrate that RERL can leverage improvements discovered by less conservative replicas, which improves task performance relative to reinforcement learning without exchange. We also observe amortized overall compute costs, which demonstrate RERL’s better overall sampling efficiency.
Chat is not available.
Successful Page Load