Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor’s Internal States
Yunho Choi ⋅ Jongwon Lim ⋅ woojin Ahn ⋅ Minjae Oh ⋅ Jeonghoon Shim ⋅ Yohan Jo
Abstract
Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models rests on variance reduction, which requires both a reliable baseline and high prompt diversity within each training batch. This is especially difficult in multi-domain training for general reasoning models, where prompts from different tasks induce highly diverse gradient signals. Existing approaches fall short in different ways: GRPO estimates its baseline as the group mean over rollouts from the same prompt, so an accurate baseline leaves fewer distinct prompts in the batch, while PPO avoids this trade-off by training a policy scale critic, roughly doubling the cost of training. We introduce $\textbf{POISE}$ ($\underline{P}$olicy $\underline{O}$ptimization with $\underline{I}$nternal $\underline{S}$tate Value $\underline{E}$stimation), a reinforcement learning algorithm that turns the model's internal states into a value model. A lightweight probe reads the signals already computed during the forward pass to predict the baseline, and is trained online alongside the policy. To preserve gradient unbiasedness, we introduce a cross-rollout construction that predicts each rollout's value from an independent rollout's internal states. On Qwen3-4B and OLMo3-7B-Instruct-DPO across a six-domain verifiable-reward corpus, POISE outperforms other RLVR baselines while maintaining lower gradient variance. Moreover, the probe matches a separate LLM-scale value model, generalizes to various tasks, and remains accurate as the policy scales. By leveraging the model's internal representations, POISE enables stable policy optimization.
Chat is not available.
Successful Page Load