Correlation-Aware Exploration for Long-Trace Reasoning
Junxing Hu ⋅ Zicheng Zhang ⋅ Zhiqian Zhang ⋅ Ai Han ⋅ Yifeng Zhang ⋅ Pengzhang Liu ⋅ Ling Li
Abstract
Reasoning models trained with Group Relative Policy Optimization (GRPO) now produce chains of thought thousands of tokens long, and recent studies reveal that such training is prone to policy collapse. Entropy regularization is the standard remedy, but it works by pushing probability mass away from the few tokens that carry it and toward the vast low-probability tail of the vocabulary, and it applies this force once for every token of the reasoning trace. This puts the bonus in direct conflict with the entropy decay that convergence requires, and makes its coefficient difficult to set in the long-trace regime. To resolve this tension, we propose Info-GRPO, an information-theoretic framework that cultivates correlation between the policy and a latent prior. Rather than raising entropy everywhere, Info-GRPO builds on the mutual information between a latent variable and the trajectory it induces: augmenting prompts with latent seeds during training lets the model explore diverse policies correlated with the prior, while the seed-conditioned entropy is guided toward convergence. Diversity is therefore carried across whole reasoning traces rather than by per-token randomness, and exploration no longer fights entropy reduction. Experiments show that Info-GRPO outperforms vanilla GRPO and entropy-regularized GRPO across diverse reasoning benchmarks. On AIME24, it improves $Avg@8$ over GRPO by 3.75\%, 1.66\%, 4.16\%, and 4.17\% on four backbones, and it reaches majority-vote consensus with up to $16\times$ less test-time compute. Code and models will be publicly released.
Chat is not available.
Successful Page Load