The Physical Cost of Context: Quality-Constrained Resource Accounting for Coding Agents
Abstract
Context policies for coding agents are usually evaluated with prompt tokens and task success, although multi-turn execution also consumes uncached prefill, KV-cache residency, time, and energy. We introduce a quality-gated physical-accounting protocol that charges failed trajectories, amortizes the retrieval index, and separates logical prompt volume from reusable prefixes. On 450 real-issue localization cases derived from SWE-rebench V2, eager evidence delivery meets a prespecified five-point non-inferiority gate. In a separate 900-cell instrumented study, it cuts model calls by 35.7\%, wall time by 29.4\%, and the GPU-energy proxy by 33.2\%; anchored history reduction provides no further saving. A 100-instance SWE-bench Verified study with mini-SWE-agent finds that last-five observation masking cuts prompt volume by 33.3\% but loses 135.9k reusable tokens per task. Uncached prefill rises by 608.5\%, the energy proxy rises by 125.6\%, and the policy misses the quality gate. Early evidence saves work by shortening search, while history deletion can increase executed work by invalidating prefix reuse.