TREEHCA: Tree-Based Hindsight Credit Assignment for Multi-Turn LLM Agents
Abstract
Reinforcement learning for multi-turn LLM agents is challenging because sparse terminal rewards provide little guidance about which intermediate decisions determined the final outcome. Tree-structured rollouts provide a natural foundation for addressing this challenge. By sharing interaction prefixes, they avoid repeatedly sampling complete trajectories and allow the rollout budget to be concentrated on selected intermediate states. Branches from the same state also explore alternative continuations, turning terminal outcomes into counterfactual, process-level learning signals. We introduce \TreeHCA, a unified framework that uses hindsight to guide both exploration and credit assignment over these tree-structured rollouts. During tree construction, \TreeHCA identifies consequential decision points where an agent begins to deviate from a successful solution and branches from them to explore corrective alternatives before errors propagate through subsequent interactions. After collecting the rollouts, \TreeHCA propagates each leaf-level advantage back through its trajectory, distributing the learning signal among intermediate decisions in proportion to their estimated contribution to the outcome. This concentrates policy updates on decisions that meaningfully affected success or failure while reducing credit assigned to incidental steps. We evaluate \TreeHCA with Qwen3-4B and Qwen3-8B on seven single-hop and multi-hop open-domain question-answering benchmarks with search tools. Under matched tree-expansion settings, \TreeHCA achieves higher mean performance than strong tree-based reinforcement-learning baselines across both model scales.