Prefix Stability for Prompt Caching in LLM Agents: When It Helps and When It Backfires
Prashant Kumar
Abstract
Iterative LLM agents (e.g., ReAct loops) re-feed a growing context on every turn. Provider-side prompt caching can amortize this by reusing a previously processed prefix at a steep discount, but only when consecutive requests share a byte-identical prefix. A natural agent design, switching the system prompt between reasoning phases, changes the prefix early and disables caching silently: no error, correct answer, full price. We formalize the prefix-reuse conditions for an agent loop, show that full-rate input-token processing collapses from $\Theta(N^2)$ to $\Theta(N)$ under an invariant, append-only prefix, and prove that a phase-switched loop can cost more than no caching at all, because write premiums are paid but never read back. We validate across four prompt architectures on two production models: the data confirm every prediction, including a measured 24.6% cost increase when caching is misapplied to a phase-switched loop, an architecture that caches nothing at the cost of no caching, and a model-specific cacheable floor that silences caching on one model at a prefix size another caches. Our contribution is the formal, application-level account of when and why violating prefix stability fails for iterative loops: a cost-complexity result, the first proof that prompt-caching misuse can invert into a net cost increase, a four-architecture taxonomy with a developer decision procedure, and an empirical validation on production systems.
Chat is not available.
Successful Page Load