Diagnosing and Stabilizing RLVR Training through Entropy Dynamics
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) is a central component of modern LLM post-training, yet the training is prone to instability and collapse. Prior work has largely studied this instability through the lens of off-policiness, but the resulting mitigation strategies do not consistently stabilize training across settings. Rather than attributing instability to a particular training–inference mismatch, we take a complementary approach that characterizes and directly controls a common signature of collapse. Across models and datasets, we find that RLVR collapse is consistently associated with a sharp increase in entropy. We use a first-order estimate of entropy change from the policy gradient and show that it closely tracks entropy dynamics during training. Directly controlling updates according to this quantity produces consistent improvements in training stability across models and datasets. We further find that the effectiveness of existing stabilization strategies is closely associated with their ability to control this entropy-gradient signal, providing a common explanation for when these interventions succeed or fail. Together, these results provide an entropy-based perspective on RLVR instability and suggest that controlling entropy growth is a key ingredient for stable training.