SIRAS: Sibling-Relative Advantage Shaping for Reinforcement Learning from Verifiable Rewards
Abstract
Reinforcement learning from verifiable rewards (RLVR) trains reasoning language models with group-relative objectives such as GRPO. However, it leaves two sources of signal underused: every token in a rollout receives the same response-level advantage, and groups in which all rollouts succeed or all fail have zero group-relative advantage and are typically discarded. We observe that same prompt rollouts are not independent reward labels but sibling attempts: they share scaffolding, diverge at uncertain decisions, and sometimes reconverge. In this paper, we propose Sibling-Relative Advantage Shaping (SIRAS), a lightweight advantage-shaping method that exploits this structure. SIRAS segments rollouts into chunks, aligns sibling chunks with banded dynamic time warping, and reshapes the flat GRPO advantage across tokens via soft-divergence weights that upweight chunks distinguishing a rollout from its siblings and downweight shared ones. This reshaping preserves each rollout’s token-averaged GRPO advantage exactly and reduces to GRPO when weights are constant. The same sibling structure produces calibrated residuals for all-correct and all-wrong groups, reclaiming training signal that group-relative baselines discard. Across six reasoning benchmarks, SIRAS yields the largest gains on AIME-style competition math, including +12.2 on AIME25 with Qwen3-8B and +23.3 on AIME26 in our backbone-scaling study, without extra rollouts, process reward models, or step annotations.