Length, Not Bit Width: KV-Cache Quantization Moves the Long-Context Retrieval Cliff on a 512 GB Workstation
Abstract
Quantizing the key-value cache is the standard response to long-context memory pressure, and it is usually evaluated on hardware where the full-precision baseline cannot run, so its cost relative to that baseline is rarely measured. We measure it. On a 512 GB unified-memory workstation, where FP16 KV at 128K context is feasible for a dense 70B model, we benchmark single-needle retrieval for Llama-3.1-70B and Qwen-2.5-72B across five needle depths and seven context lengths from 4K to 128K, at FP16, int8 and int4 KV precision, for 190 logged runs. Three findings. First, retrieval fails as a cliff in context length rather than a decay across depth: below a model-specific length every depth succeeds, above it nearly every depth fails, and pooled depth marginals span only 55-66%. Second, affine KV quantization moves the cliff inward. Qwen retrieves at every depth through 96K at FP16 under lenient grading, 64K under strict grading, and only through 32K once the cache is quantized, a two- to threefold reduction in usable context; Llama collapses from partial retrieval at 64K to none. Bit width does not matter: int8 and int4 give identical pass/fail patterns on Llama and differ by one cell on Qwen. Third, in the implementation tested the compression does not pay for itself on the other side of the trade. Peak memory was higher under quantization than at FP16 in all 24 model, length and bit-width configurations, by 1 to 28 GB, and prefill throughput roughly halved at 128K, so the penalty is smallest where compression is unnecessary and largest where it is the point. We argue that KV compression methods should report where the retrieval cliff sits against a full-precision baseline, not a mean over a context sweep, and we release the harness and all per-run logs.