Progress Is Not Prognosis: Local State and Eventual Outcome in Long-Horizon Agents
Abdelmageed Abdelmagid ⋅ Justin S Sidharta ⋅ Ryan Xing ⋅ jeff sun ⋅ Jonas Rohweder
Abstract
Language models have become widely used as agents for long-horizon tasks [1]. This makes it critical to understand whether representations learned from short-horizon settings preserve their meaning during long-horizon execution, particularly for diagnosing failures and overseeing agent behavior. Recent work has identified a value axis in single-turn settings that tracks a model's internal expectation of success for its current reasoning [2]. We ask whether this representation keeps its meaning when a model operates as a long-horizon agent. We reconstruct the value axis of Jiang et al. [2] on Qwen3-32B and apply it without refitting to 500 SWE-bench Verified trajectories. The frozen axis responds reliably to local evidence of failure, decreasing after environment errors and recovering over the following steps. It also separates resolved from unresolved trajectories when pooled across tasks (AUROC $0.747$). However, this separation is largely explained by task difficulty and disappears once task identity is held fixed. Thus, the axis is sensitive to the current state of an execution, but does not reliably encode its eventual outcome. We then identify a separate direction that predicts eventual resolution when task identity is held fixed, correctly ranking the resolved trajectory above the unresolved trajectory in 80.3\% of held-out pairs. The outcome direction does not respond to local errors, and removing the value-axis component leaves its performance unchanged. Overall, these results suggest that long-horizon agents represent local execution state and eventual outcome through distinct signals.
Chat is not available.
Successful Page Load