What Makes Two States the Same? Value-Aware State Matching for Multi-Turn Agent Credit Assignment
Abstract
Training multi-turn LLM agents with reinforcement learning requires distributing a single episode reward across many interaction steps. Recent group-based methods tackle this credit assignment challenge by finding steps where the agent visits the "same" state across different rollouts and comparing their outcomes. However, determining when two states are "the same" remains an open problem: existing methods compare raw text observations via exact string matching or bag-of-words similarity, both of which ignore whether state differences actually matter for future returns. We show that this is a state abstraction problem with a precise bias-variance tradeoff: matching on value-irrelevant features creates unnecessarily small groups (high variance), while ignoring value-relevant features merges states that should be distinguished (high bias). Guided by this analysis, we propose Structured State Abstraction (SSA), a step-level credit assignment method that decomposes observations into typed fields using environment-provided metadata and learns per-field importance weights from return variance. Rather than hard grouping, SSA uses a kernel-smoothed advantage estimator where all states contribute to each step's baseline proportionally to their structured similarity, yielding lower MSE than threshold-based grouping. Experiments on ALFWorld and WebShop show consistent improvements over GiGPO and HGPO across both benchmarks and model scales; notably, kernel smoothing eliminates the singleton problem that causes existing methods to lose step-level credit for 30-36% of steps. Our implementation is provided in the supplementary materials and will be open-sourced upon acceptance.