SeRPO: Segment-level Rubric-based Policy Optimization for LLM Agents
Abstract
Reinforcement learning for long-horizon LLM agents is bottlenecked by its reward signal: a binary, trajectory-level reward is almost always zero and gives no gradient when successful rollouts are rare. Prior remedies either propagate the outcome to individual steps or replace the binary signal with richer rubric scores, but step-level credit still rests on the outcome and rubric scores are applied to the whole trajectory. We propose SeRPO (Segment-level Rubric-based Policy Optimization), in which a single frozen-LLM call jointly partitions a trajectory into subgoal-aligned segments and scores each segment's contribution, feeding the segment-level advantages to GRPO with no separately trained reward model. On AppWorld, SeRPO outperforms outcome-reward baselines at the trajectory level and at the step level. Notably, step-level credit raises per-task completion but leaves scenario completion unchanged from the trajectory baselines, whereas SeRPO improves both. And because the segment score is independent of the outcome, SeRPO yields a learning signal even when every rollout fails: trained on failed trajectories alone, where binary and step-level rewards give no gradient at all, it still improves over base.