Causal Decomposition of LLM Agent Failure via Oracle-State Intervention
Abstract
LLM agents degrade as trajectories lengthen, motivating current research on context management methods that compress, filter, or restructure history. All of these methods share the untested assumption that context degradation is the primary bottleneck of agentic systems. To test this hypothesis, we introduce an oracle-state intervention framework that decomposes agent failure modes, isolates context as a contributing variable by replacing a failing agent's accumulated history with ground-truth structured state, and measures recovery. Across 420 failed WebArena trajectories with two models (Qwen3.5-27B and GPT-5.2), we find that context management is effective but not in the way commonly assumed in recent literature. In the default WebArena Qwen and GPT configurations, replacing context after failure recovers under 10\% of failed trajectories. Additionally, independent human annotators attribute only 11.7\% of failures to context-caused errors. The remaining 58–74\% of failures persist even after both context and environment are restored, and the convergence between intervention-based and human-attribution estimates suggests that context availability explains only a minority of failures. Yet continuous oracle state throughout the trajectory improves success by 19-38\%, and the gap between oracle-maintained and self-generated state widens from 8\% to 12.5\% as trajectories lengthen. This controlled intervention framework reveals that post hoc context repair rarely rescues failed agents, while continuously maintained privileged state improves long-horizon execution. Moreover, our findings suggest that progress on agentic systems will require improving planning and execution capabilities alongside context engineering, rather than treating context quality as the dominant bottleneck.