From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space
Yuqiao Tan ⋅ Minzheng Wang ⋅ Bo Liu ⋅ Zichen Liu ⋅ Shizhu He ⋅ Jun Zhao ⋅ Kang Liu
Abstract
Reinforcement learning with verifiable rewards (RLVR) enhances LLM reasoning by optimizing the conditional distribution $P(y|x)$, but its gains remain constrained by the base model's output distribution. In this paper, we introduce PreRL (Pre-train Space RL), which optimizes the unconditional distribution $P(y)$ through reward-driven online updates. Unlike conventional pre-training on static corpora, PreRL actively aligns the pre-train space with reasoning tasks, enhancing reasoning ability while preserving broad exploration capacity. We theoretically and empirically show that gradients of $\log P(y)$ and $\log P(y|x)$ are strongly aligned, making PreRL a viable surrogate for standard RL. We further identify Negative Sample Reinforcement (NSR) as the key mechanism in PreRL: NSR-PreRL prunes incorrect reasoning paths and elicits reflective behaviors. Leveraging these insights, we propose Dual Space RL (DSRL), a Policy Reincarnation strategy that initializes models with NSR-PreRL to expand the reasoning horizon before transitioning to standard RL for fine-grained optimization. Extensive experiments demonstrate that DSRL consistently outperforms strong baselines, effectively steering the policy toward a refined correct reasoning subspace.
Chat is not available.
Successful Page Load