AgentKVShift: Efficient KV Cache Reuse for Agentic Memory Systems
Abstract
Memory-augmented LLM agents have gained substantial attention recently for their ability to maintain context across hundreds of interactions through agentic memory systems that actively curate retrieved content with LLM-generated metadata such as summaries, keywords, and tags. However, from an inference cost standpoint, every retrieval triggers a full re-encoding of these structured memory units into Key-Value (KV) states, which dominates prefill latency. Existing training-free KV reuse methods mitigate this by selectively recomputing a small fraction of tokens, but were designed for RAG-style raw passages and degrade substantially on structured agentic memories. In this work, we present AgentKVShift, a training-free, probe-guided KV residual correction method that operates per retrieved memory unit. One of the crucial insights we demonstrate is that the per-memory KV reuse residual decomposes into a shared memory-level offset plus small token-wise fluctuations. Estimating this offset from a small probe set allows us to correct every reused token in the memory unit by a single weighted correction. Unlike prior reuse methods which decide which tokens to recompute and leave the rest of the cache stale, AgentKVShift also corrects the tokens it does not recompute, turning the refresh budget into useful signal across the entire chunk. Through extensive experiments across four open source LLMs spanning 3B to 32B parameters and two long-horizon agentic memory benchmarks covering long-term dialogue and agentic applications, we show that AgentKVShift achieves near full recompute performance while refreshing only 10-30% of the cache, outperforming existing baselines at the same recompute ratio. AgentKVShift requires up to 5x lower recompute to reach this near-full performance, which prior reuse methods only attain at 45-55% refresh. By operating in this lower recompute regime, AgentKVShift delivers prefill speedups of 2-3.5x over no-KV-reuse on a single A100 GPU. Lastly, AgentKVShift orthogonally composes with KV cache quantization, retaining over 2x the F1 of prior reuse methods under aggressive 2- and 4-bit settings, making it a simple, yet effective choice for serving long-horizon agentic memory workloads.