When Does Knowing the State Help? Diagnosing Process vs. Outcome Reward Design
Abstract
Process reward models (PRMs) dramatically outperform outcome reward models (ORMs) for math reasoning and agent tasks, yet provide no benefit for chat and QA. No theory explains why, and practitioners choose between them by expensive trial and error. We introduce state-lift, a diagnostic that predicts whether a PRM is worth training for a new domain — in minutes, from ~500 labeled steps, before committing to expensive reward model training. State-lift measures how much step quality depends on trajectory state versus action text alone, and an accompanying effective gap distinguishes two regimes: state-noise-dominant (PRM essential) and action-gap-dominant (ORM sufficient). Across ten domains and eleven published PRM-ORM comparisons, state-lift correctly predicts PRM advantage, and a minutes-scale linear estimate matches the outcome of full Llama-3.1-8B reward model training — including CaSiNo negotiation (+0.055 lift at SL = 0.22). The accompanying two-branch decision protocol turns state-lift into an actionable PRM-vs-ORM workflow that runs in minutes on a single CPU.