Where Should a Stuck Agent Spend Its Next Tokens? Causal Attribution of Inference-Time Resources
Saiyue Lyu ⋅ Jerrold Huang ⋅ Yunbei Zhang ⋅ Janet Wang ⋅ Zhenting Wang ⋅ Yingqiang Ge ⋅ Haibo Zhang
Abstract
Inference-time performance is often plotted against token budget. A coding agent facing a failure, however, can seek more computation, information, or feedback. Which resource helps, and does its value depend on what is already available? We introduce Banzhaf Resource Attribution (BRA), a controlled factorial framework for estimating each resource's average marginal contribution to repair and its interactions with others. BRA holds the failed state and final revision protocol fixed, attributing differences in held-out repair utility to controlled resource interventions. We evaluate six models on 194 function-level and 126 repository-level failures. Under the tested conditions, audited specifications improve repair utility for every model, while public-test feedback provides small gains and retrieved-code packets yield approximately zero gain. Additional computation becomes less valuable across the evaluated capability ordering, whereas audited information remains beneficial. Pairwise resource interactions are negative or near zero, while fix-location and required-behaviour information complement one another. Yet three of four models in a resource-choice probe select feedback in $95$--$100\%$ of cases, obtaining lower gains than uniform random selection. BRA thus reveals a gap between agents' resource preferences and measured repair benefits.
Chat is not available.
Successful Page Load