SPRM: From Cooperative Games to Marginal-Contribution Process Reward Modeling
Abstract
Process reward models provide step-level feedback for reasoning, but reliable process supervision remains difficult to obtain: human annotations are costly, Monte Carlo rollouts require large sampling budgets, and LLM judges can introduce prompt-sensitive biases. We introduce SPRM, an outcome-only framework that constructs process rewards directly from terminal correctness signals. SPRM views a reasoning trajectory as a temporally ordered cooperative game, where reasoning steps act as participants and the final outcome reward is the shared payoff. Following this allocation view, SPRM learns a differentiable trajectory value function from outcome rewards and distributes the payoff to individual steps through Aumann-Shapley-style Step Value Integration. We further analyze its Shapley properties under the temporal and causal structure of reasoning trajectories. To make the resulting supervision robust, SPRM incorporates causal consistency and estimates a cross-trajectory inherent score that captures each step's context-independent value from semantically similar steps. These signals train a dual-head process reward model that jointly represents trajectory-specific marginal credit and cross-trajectory inherent step value. Across ProcessBench, PRMBench, and Best-of-N reranking on AceMath-RewardBench, SPRM improves first-error localization, marginal-contribution-sensitive evaluation, and test-time reasoning selection over existing open PRMs trained from human, rollout, curated, or outcome-broadcast supervision. These results suggest that marginal-contribution-based credit redistribution offers an effective and scalable route to process supervision from outcome-only data.