OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in Policy Optimization
Abstract
Reinforcement learning methods for LLM reasoning typically assign a uniform advantage to every token in a trajectory, diluting the learning signal at pivotal reasoning steps and injecting noise at uninformative positions. Critic-free alternatives such as anchored distillation derive per-token signals from oracle-conditioned likelihood ratios, but apply the signal independently at each position. We propose Oracle-Prompted Policy Optimization (\n), which recovers exact token-level advantages through a Bayesian value recursion. Conditioning the model on the ground-truth answer during training yields per-token likelihood ratios that measure how much the answer revises the model's prediction; the ratios accumulate via Bayesian updating into a running estimate of the success probability at every position, from which the token-level advantage follows in closed form. The advantage factors as a product of a state weight, peaking where the outcome is most uncertain and vanishing where success or failure is already determined, and the per-token oracle evidence familiar from on-policy distillation, with no tunable weighting hyperparameter. The framework supports two estimator choices: a self-oracle mode in which the policy model provides the estimate through one additional forward pass, and a teacher-oracle mode in which a stronger model serves as the estimator. Experiments on two base models across seven reasoning benchmarks spanning mathematics, science, and code demonstrate consistent improvements over GRPO, DAPO, and SDPO, with the largest gains on competition-level tasks where reasoning chains are longest.