Matched Budgets Need Halt Rates: Auditing Agent Architecture Comparisons
Abstract
Matched token budgets are the standard control when comparing LLM agent architectures, but a success rate under a shared ceiling cannot distinguish an architecture that performed worse from one that never finished its control process. We argue that four quantities must be reported together: success, budget-halt rate, per-component resource share, and component control authority; and we show what each one changes. On a deterministic security-incident simulator we run 810 live episodes (90 incidents × 3 trials × 3 architectures) at a shared 96k ceiling, plus 72 live episodes on a second, stronger model. Three findings follow. First, hierarchy is the wrong unit of analysis: a planner– worker scaffold matches flat ReAct’s token cost and beats it by 7.4 points, while adding a verifier doubles median consumption (10,696 to 22,315 tokens) for −2.6 points. Second, the verifier is not inert: it parses at 99.8% and changes control flow in 4.7% of 2,030 invocations, yet it consumes 48.5% of the hierarchy’s ledger, so cost share and control authority must be measured separately before an ablation means anything. Third, the architecture verdict is capability-conditional: on a model with a 35.2% flat baseline the full hierarchy ends +4.8 points ahead at zero halt, while on a model with a 91.7% flat baseline it ends −11.1 points behind and is still 8% halted. We also show, by retrospective truncation of the recorded trajectories, that at a 24k ceiling 45% of hierarchy episodes are budget-terminated against 6% of flat episodes, which is enough to invert the apparent ranking. We state the assumptions this truncation requires, what it does and does not establish, and package the four measurements as a reporting protocol recoverable from logs a deployed agent already produces.