Exploration-Preserving Policy Optimization
Abstract
Reinforcement learning with verifiable rewards improves reasoning, yet group-relative objectives give equal credit to samples with equal rewards. Frequent response modes therefore accumulate more update mass than rare modes, narrowing correct support and reducing the value of repeated sampling. We introduce Exploration-Preserving Policy Optimization (ExPPO), a lightweight advantage-shaping layer that combines rollout-policy surprisal with prompt pass rate to rebalance positive and negative credit across response modes. ExPPO preserves reward polarity, bounds the correction, and approximately preserves prompt-level absolute advantage mass. A local KL-regularized analysis characterizes its mode shifts and conditions for entropy and target-utility improvement. Across model families and scales on reasoning tasks, ExPPO delivers stronger in-domain adaptation and out-of-domain generalization than standard group-relative and entropy-regularized baselines, achieving the best final coverage averages for every evaluated backbone. Large-budget sampling and entropy dynamics show broader correct support without indiscriminate uncertainty, while a controlled multi-answer evaluation demonstrates a broader conditional distribution over correct solution modes than GRPO. ExPPO thus improves reasoning coverage by allocating learning signal according to response rarity.