Step-wise Rubric Rewards for LLM Reasoning
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve the reasoning capabilities of large language models, but its reward is derived only from the correctness of the final answer and provides no supervision over intermediate reasoning steps. Recent rubric-based methods, such as Rubrics as Rewards (RaR), introduce finer-grained supervision by scoring rollouts against a structured set of evaluation criteria. However, the resulting rubric scores are still aggregated into a single scalar that is applied to the entire response, leading to three structural weaknesses, namely loss of the multi-criterion rubric structure, uniform supervision of correct and incorrect reasoning steps, and reward hacking in the trained model through unbounded self-correction. On a sample of 1{,}000 problems, we find that 18.2\% of steps within answer-correct responses are themselves wrong yet positively rewarded, while 49.9\% of steps within answer-incorrect responses are in fact correct yet penalized. Therefore, we introduce \textbf{Step-wise Rubrics as Rewards (SRaR)}, an RLVR framework that (i) uses an LLM judge to attribute each rubric item to a specific reasoning step, (ii) normalizes the per-step rubric scores across rollouts so that only steps whose quality varies produce a learning signal, and (iii) combines the resulting per-step reward with the standard outcome reward through a decoupled advantage estimator that keeps the outcome-driven baseline stable. To support training, we further build a 16K-problem rubric dataset by contrastively distilling rubric items from correct and flawed reasoning paths sampled from a strong model and verified against the ground-truth answer. Across six mathematical reasoning benchmarks spanning multiple difficulty levels, SRaR improves the average accuracy over RaR by \textbf{3.57} points on Qwen3-8B-Non-Thinking and by \textbf{2.75} points on Qwen3-32B-Non-Thinking, raises the Faithful Reasoning Rate on AIME~2025 from 34.5\% to 46.7\% (every correct-answer trajectory of SRaR uses entirely correct reasoning steps), and reduces the rate of self-correction looping from 48.1\% to 26.5\%.