Twin-Turbo Quantizer: Accurate Low-Bit Transformer Attention via Temporal Smoothness
Lucas P Goetz
Abstract
The Twin-Turbo quantizer decorrelates the key-value (KV) cache along the temporal axis to achieve practical, near-lossless FP4 quantization. The proposed quantization scheme factorizes blocks of the KV cache into a rank-1 matrix and dense residual matrices, enabling split bitwidth allocation to shared and individual token directions. Our experimental findings show that the Twin-Turbo quantizer can achieve accuracy comparable to 16-bit baselines while requiring no calibration data and remaining agnostic to the underlying vector quantizer for decode workflows. Furthermore, the quantizer provides a single interpretable parameter to trade between accuracy and efficiency: the quantization block size.
Chat is not available.
Successful Page Load