Where Should the Marginal Dollar Go? Cost-Performance Frontiers and Reflection-Guided Verification for GAIA Agents
Abstract
Public agent benchmarks rank systems by accuracy as if compute were free. We build a complete per-task dollar-attribution pipeline for a tool-using agent on the GAIA validation split (165 tasks, public gold; the leaderboard 93 percent figure is a different held-out test split and not comparable to any number in this paper) and trace a cost-performance frontier across five configurations spanning 13 times in spend. Four form a clean Pareto frontier: GPT-5.6 Luna (0.020 dollars per task, 63.0 percent), DeepSeek V4 Flash (0.039, 70.9 percent), a Flash-solver plus Pro-0813-verifier cascade (0.107, 74.5 percent), and Pro-0813 in both roles (0.265, 78.8 percent). Qwen 3.7 Plus is strictly dominated, costing 68 percent more per task than Flash while solving three fewer of 165 tasks. The endpoints separate significantly (plus 7.9 pp, exact McNemar p=0.041, bootstrap 95 percent CI 0.6 to 14.5 pp), and the marginal dollar buys unevenly: verification costs 1.32 dollars per additional correct answer at GAIA Level 2 versus 2.39 to 3.53 dollars at Levels 1 and 3. Sticker price is a poor guide throughout. Luna and Flash list within 3 percent per blended token yet differ by a factor of 2 in measured spend, because token intensity rather than price drives agent cost, a pattern confirmed on GPT-5.6 Terra/Sol diagnostic canaries where a 10 times list premium yields only a 2 times run-cost premium. A five-step evidence-reflection loop attributes the gain mechanistically: the verifier is strictly additive (17 fixed, 1 broken) while the harness answer-selector disproportionately overrides the most expensive tier's own correct answers. We also report where this stops working: the verifier rubber-stamps an already-wrong solver answer in 28 of 29 disagreements, and three pre-registered redesigns each fail their gate. An external audit flags 16.4 percent of GAIA validation as defective, capping honest accuracy near 89 percent. Finally, an ensemble over pooled solver trajectories reaches 143 of 165 strict (86.7 percent), eight points above the best single configuration. We release the full per-task evidence corpus, reflection traces, token-intensity tables, and cost-attribution pipeline.