Prefill, Decode, and Context: A Measured Energy Law for Agentic LLM Serving
Abstract
Deploying a multi-agent LLM system (MAS) multiplies LLM calls across agents, tool-calling steps, and communication rounds, yet the energy this actually costs is almost never measured: prior work prices compute in tokens, calls, or dollars, and mostly on closed API models where energy is not measurable by the user at all. We instrument every LLM call and tool execution of a locally served MAS grid (four topologies, four agentic benchmarks, primarily a 9B model) with GPU hardware energy counters. From these records, we fit a three-term physical model of per-call energy: prefill, a decode term, and a context term for re-reading the key-value (KV) cache, which achieves R² = 0.961 and is preferred over token-count-only baselines. On our prefix-cached, batch-1 grid, a decode token costs ≈335× a prefill token. This asymmetry depends heavily on the serving regime: in dedicated serving-occupancy experiments (batch sizes 1–128), it falls from ≈115× at uncached batch 1 to a ≈3.3× hardware floor under production-style batching. The KV-context term, by contrast, does not amortize with batch size at all. We turn this into a parameter-free law of how badly single-rate token or dollar accounting misprices agentic energy: a 20× spread across cells under serial serving, still 1.4× at the production floor. The same physical form, with interpretable coefficient shifts, transfers across a dense 9B model, a 35B mixture-of-experts model, and a quantized 31B dense model. Our results give resource-aware agentic-AI deployments a validated, hardware-grounded accounting of what an LLM call costs, and where token- or dollar-based proxies for that cost break down.