RankVQ: Low-Rank Parameterized Commutative Vector Quantization for KV Cache Compression
Jianglin Zhou ⋅ Huaming Wu ⋅ Huijun Tang
Abstract
The growing context lengths of large language models (LLMs) have made key-value (KV) cache memory a critical bottleneck during inference. Vector quantization (VQ) with commutative codebooks provides an effective Rotary Position Embedding (RoPE)-aware solution by enabling reconstruction-free attention over quantized key indices. However, existing commutative VQ methods still parameterize scalar codebook coefficients across RoPE subspaces and codewords largely independently, leaving parameter-side redundancy underexploited. We propose $\textbf{RankVQ}$, a low-rank parameterized commutative vector quantization method for KV cache compression. Instead of factorizing the deployed codebooks, RankVQ factorizes the underlying scalar coefficient matrices used to construct key and value codebooks. This design preserves the pseudo-symmetric structure required for RoPE commutativity on the key side, while improving the quality and compactness of offline codebook construction. The low-rank factors are discarded after codebook construction, so RankVQ keeps the same online inference form as standard commutative VQ and introduces no additional decoding overhead. Experiments on LongBench, InfiniteBench, and GSM8K show that RankVQ achieves a stronger accuracy--compression trade-off than directly comparable KV quantization baselines, with particularly clear gains in the 1-bit regime.
Chat is not available.
Successful Page Load