OAMP-TQ: Ultra-Low-Bit TurboQuant KV Cache Compression for Small Language Models
Abstract
Small Language Models (SLMs) are increasingly considered for efficient deployment, including agentic applications that require long-context processing. However, the growing Key-Value (KV) cache remains a major memory bottleneck. While vector quantization methods such as TurboQuant (TQ) achieve near-optimal distortion rates, reducing KV cache to uniform 2-bit precision can substantially degrade accuracy due to outlier activations. We propose OAMP-TQ, an outlier-aware mixed-precision TurboQuant method for aggressive KV cache compression, evaluated on LLaMA, Mistral, and Ministral 7–8B instruction-tuned models. We introduce two strategies: (1) One-Level (1L) OAMP-TQ, which quantizes inlier channels to 2-bit while preserving 5–10\% of outlier channels in FP16; and Two-Level (2L) OAMP-TQ, a fully mixed-integer scheme applying 2-bit TurboQuant to inliers and 4-, 6-, or 8-bit TurboQuant to outliers. We evaluate both against unquantized KV cache and 4-bit TurboQuant on Needle-in-a-Haystack (NIAH) and LongBench (LB). 1L OAMP-TQ performs comparably to 4-bit TQ on LB but provides only 15\% memory reduction and reduces NIAH recall from 98.8\% to 88.6\%. In contrast, 2L OAMP-TQ with 2-bit inliers and 4-bit outliers reduces KV cache memory by 45\% over the 4-bit baseline with comparable quantization time while maintaining strong NIAH accuracy. These results show that outlier-aware mixed-precision quantization enables aggressive KV cache compression while preserving long-context accuracy, supporting memory-efficient SLM deployment in long-context and agentic settings.