Skill-Coupled Policy Optimization with Calibrated Group-Wise Advantage Estimation
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models on reasoning tasks. To avoid the complexity of training explicit critics, algorithms like GRPO estimate advantages directly from sparse on-policy rollouts. However, the efficacy of this approach relies heavily on the accuracy of the baseline estimator. In this work, we argue that standard methods, whether relying on prompt-local averages or shrinking toward a single global mean, are statistically mis-specified for reasoning domains. By enforcing a uniform prior, these methods ignore the latent skill structure of tasks, introducing estimation bias that conflates dissimilar problems. To address this, we propose Skill-Coupled Policy Optimization (SCPO), a structured baseline estimation framework that refines the shrinkage prior by leveraging task correlations. Instead of a generic global target, SCPO performs shrinkage toward skill-coupled group targets, thereby aligning the baseline with the local difficulty landscape. To further stabilize these targets under limited on-policy observations, SCPO incorporates a lightweight history tracker. Experiments across diverse models and tasks show that SCPO consistently improves training stability and performance over both prompt-level and global shrinkage baselines.