Where Reliability Comes From: Stagewise Evaluation of Hierarchical LLM Agents
Abstract
Final-task accuracy does not reveal whether a hierarchical agent succeeds through correct upstream work or downstream error repair. We study this distinction in a two-stage worker–parent pipeline, crossing worker execution feedback with advice-only versus reusable-artifact handoff while holding parent configuration and final authority fixed. Scoring both the worker candidate and final submission yields an exact realized accounting of upstream success, downstream repair, fallback, and damage. On 25 SpreadsheetBench Verified tasks with Claude Sonnet 4.5, the preregistered feedback-by-handoff interaction is unsupported (δY = 0.00, 95% CI [−0.16, 0.16]). Intermediate outcomes nevertheless reveal a marked redistribution: production-candidate success rises from 17/25 to 21/25, observed downstream repairs fall from five to zero, and final success changes from 22/25 to 21/25. The five baseline failures repaired downstream are exactly the five that become upstream successes under worker feedback. This exploratory alignment suggests overlapping execution-visible error coverage across stages. A preregistered paired replay of 50 fixed candidates, a prospectively frozen eight-issue SWE-bench study, and a separately calibrated GPT-5.6 Luna policy stack reveal additional recovery regimes without causally identifying the mechanism. Our contribution is a stagewise evaluation approach and empirical evidence that similar endpoint performance can conceal different sources of reliability. Hierarchical-agent evaluation should therefore report the residual failures each stage resolves, alongside final accuracy and verification activity.