Intra-Trajectory Consistency for Reward Modeling
Abstract
Reward models are critical for improving large language models (LLMs), particularly in reinforcement learning from human feedback (RLHF) and inference-time verification. Due to the prohibitive cost of fine-grained annotations, current reward models typically learn from holistic response scores to determine outcome rewards. However, this coarse-grained supervision makes it difficult for the reward model to identify which specific components within a response trajectory truly correlate with the final score, leading to poor generalization to unseen responses. In this paper, we introduce an intra-trajectory consistency regularization to propagate coarse, response-level supervision into fine-grained learning signals. Inspired by a Bayesian framework, the proposed method implements a heuristic principle: the rewards of adjacent generation processes are regularized toward consistency when the connecting token has a higher generation probability, thereby propagating dependency signals across processes. We apply the proposed regularization to the outcome reward model, improving its performance on RewardBench. Furthermore, we demonstrate that the reward model trained with the proposed regularization yields better DPO-aligned and PPO-aligned policies, and achieves superior best-of-N inference-time verification results.