KV Cache Drift Across Handoffs in Long-Horizon Agents
Abstract
During handoff, a long-horizon agent often rebuilds its prompt from text processed at an earlier stage. Reusing the existing key–value (KV) cache could reduce repeated prefill; however, changes in the preceding context can change cached states. We compare sender states with states computed freshly on the receiver prompt, then measure how using the sender states changes next-token predictions. These comparisons use 35 recorded handoffs from one SWE-bench coding-agent family, with 35K–80K-token sender histories and 3.4K–25.1K-token receiver prompts. For matched tokens, we measure the smallest fraction whose exclusion brings the remaining tokens' mean normalized key error within a chosen tolerance. At a tolerance fixed from a cross-model mapper's held-out error, the median same-model removal fraction is zero. At a stricter tolerance of 0.03, it rises to 53% of matched tokens. The zero median at the reference also appears for Qwen3-4B and SmolLM3-3B on the same long handoffs and 25 shorter handoffs. The tested cross-model linear mapper shows much greater disagreement, with a median removal fraction of about 96% on the long handoffs. When the receiver reads reused states, its predictions agree with fresh-cache predictions at a median rate of 90.20% across handoffs, with a median per-handoff mean KL divergence of 0.1424 nats. Both runs receive the same recorded continuation, keeping their token histories identical. Random errors matched to the magnitude of reuse errors produce larger prediction changes, showing that equal-sized cache errors can affect predictions differently. Meeting the representation reference therefore does not ensure unchanged predictions. Evaluating reuse requires checking both cache differences and their effect on predictions. Whether reuse preserves task success or delivers practical savings remains open. Code: https://github.com/hossainpazooki/linear-ceiling.