How Do AI Agents Spend Your Money? Analyzing and Predicting Token Cost in Agentic Coding Tasks
Abstract
Wide adoption of AI agents in complex human workflows drives rapid growth of LLM token consumption. We use “token consumption” to refer to both input and output tokens used by LLM agents. When agents are deployed on tasks that can require a large amount of tokens, three questions naturally arise: (1) Where do AI agents spend the tokens? (2) What models are more token efficient? and (3) Can LLMs anticipate the token usage before task execution? In this paper, we present the first quantitative study of token consumption patterns in agentic coding. We analyze trajectories from eight frontier LLMs on a widely used coding benchmark (SWE-Bench) and evaluate models’ ability to predict their own token costs before task execution. We find that: (1) agentic tasks are uniquely expensive: they consume substantially more tokens and are correspondingly more expensive than code reasoning and code chat, with input tokens being the key driver of the overall cost, instead of output tokens. (2) Token usage is highly variable and inherently stochastic: runs on the same task can differ by up to 30× in total tokens, and higher token usage does not translate to higher accuracy; instead, accuracy often peaks at intermediate cost and degrades at higher cost. (3) Model-to-model token efficiency is governed more by model differences than by human-labeled task difficulty, and difficulty labels only weakly align with actual resource expenditure. (4) Finally, frontier models fail to accurately predict their token usage (with weak-to-moderate correlations, up to 0.39), and systematically underestimate the real token costs. Our study reveals important insights regarding the economics of AI agents and can inspire new studies in this direction.