Breaking the Role: At What Point Does Compressing a Shared KV Cache Degrade Agent Capabilities?
Pranav Sathu ⋅ Rishi Khare
Abstract
Multi-agent LLM pipelines increasingly hand off KV caches instead of text; however, existing analyses of these pipelines transfer caches uncompressed and only grade the agent that owns the cache. We present a controlled study of compressing a shared reasoning cache to evaluate whether a receiving agent still performs its role. A Qwen3-8B solver's 4.7k-token thinking trace (595\,MB in fp16) is inherited by a verifier, across a grid of (SnapKV, R-KV, StreamingLLM) $\times$ (handoff-time, decode-time) $\times$ (HQQ 2/4/8-bit), with memory reported as measured packed bytes. We grade the verifier's functionality through three metrics: how often it corrects the solver's answer, how often it breaks a correct answer, and its net value over running the solver alone. Pruning to 2048 tokens plus 4-bit quantization retains 97\% of uncompressed accuracy at $7.1\times$ less memory (0.791 vs.\ 0.813), with the correction rate intact (0.618 vs.\ 0.632). Under heavier compression, aggregate accuracy conceals asymmetric damage: at 2-bit the verifier keeps 55\% of its accuracy but only 33\% of its correction rate, begins breaking correct answers (0.000 to 0.303), and its net value-add turns negative, so the pipeline becomes worse than the solver alone. For the AIME25 dataset, quantization at 4 bits or greater matches or improves upon the uncompressed cache but every eviction condition falls below it. On long, hard reasoning traces you can afford to store everything imprecisely, but not to discard anything.
Chat is not available.
Successful Page Load