SpectralKV: Redundancy-Aware KV Cache Compression via Spectral Coreset Selection
Fangming Zhao ⋅ Fulun Ye ⋅ Xiaofei Yue ⋅ Ziming Zhao ⋅ Yu Peng ⋅ Junyu Chen ⋅ Tingting Li
Abstract
The key-value (KV) cache has become a dominant memory and bandwidth bottleneck for serving long-context large language models, motivating a growing body of work on KV cache compression. Most existing methods follow a common recipe: a uniform per-layer budget and attention top-$k$ token selection, but this recipe is \emph{data-oblivious} (ignoring cross-layer variation in attention concentration) and \emph{redundancy-blind} (retaining near-duplicate key--value entries). We propose \textbf{SpectralKV}, a training-free framework that addresses both issues jointly. Across layers, a lightweight concentration-based criterion redistributes a fixed global budget based on the entropy of observation-window attention, giving more slots to layers whose attention is more concentrated. Within each layer, a spectral coreset selector treats keys and values as geometric objects in a metric-aligned joint feature space and selects tokens via residual pivoting, suppressing near-duplicates while preserving geometrically distinct directions. On Llama-3.1-8B and Qwen3-8B, SpectralKV consistently outperforms seven strong baselines on 15 LongBench tasks and PG-19 perplexity at 16K and 32K contexts, with the largest gains at aggressive compression ratios down to $6.25\%$ keep. And a static-budget variant reduces peak prefill VRAM by up to $44\%$ with negligible quality loss. Code and data are available at \url{https://anonymous.4open.science/r/spectralKV-2FF7}.
Chat is not available.
Successful Page Load