The Long-Horizon Task Mirage? Diagnosing Why Agentic Systems Break
Abstract
Long-horizon agent reliability is increasingly important given the broad range of real-world agent applications. However, existing benchmarks lack either tasks that adequately control horizon depth or metrics to diagnose horizon-dependent failures. Across four interactive domains (web, OS, database, and embodied), we observe a common pattern: success does not decay exponentially, as would follow from the compounding of independently solvable subgoals, but collapses sharply at a point within a handful of actions as we increase the number of interdependent subgoals that must be chained. We address this gap with \texttt{HORIZON}, a cross-domain diagnostic benchmark that treats horizon as a controlled variable, and develop a seven-category failure taxonomy by analyzing failed trajectories with human experts. We further propose a trajectory-grounded LLM-as-a-Judge pipeline for scalable, reproducible failure attribution, and validate it against human annotation with high agreement. Our findings show that scaling the model alone is not sufficient: progress requires design-level advances such as execution-time verification and memory management. Above all, our work gives future agent systems a clear diagnostic map: a shared vocabulary of failure modes and metrics for pinpointing why long-horizon execution breaks down, and a concrete roadmap for future agentic research.