Parameter Exploration for RLVR via Variational Learning
Abstract
Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that encouraging exploration during LLM reinforcement learning can improve downstream performance. However, methods for controlling exploration often rely on heuristics like clipping or temperature scaling that are unable to fundamentally change token-level distributions, which might limit exploration. Here, we explore controlling exploration via variational learning, where a distribution over neural network parameters is learned via noisy optimization, with parameters sampled from an approximate posterior. We introduce Perturbed Parameter Policy Optimization (3PO), where the amount of noise that is added to model parameters functions as an additional control lever that can be tuned for exploration. We show that using multiple parameter samples within a batch performs better than using a single sample. When using multiple parameter samples in group-based methods like GRPO, we find that calculating advantages across rollouts from all parameter samples performs best. We call this Chunked Perturbed Parameter Policy Optimization (C3PO). We show empirically on a range of math reasoning benchmarks and LLMs that using multiple parameter samples can improve downstream performance, especially on harder benchmarks like AIME. Overall, our work presents evidence that parameter-space exploration can improve LLM reinforcement learning.