TACT-KV: Tri-Axis Cosine Transform for Compressing Volumetric KV Caches in Medical VLMs
Chongyu Qu ⋅ Ritchie Zhao ⋅ Yufan He ⋅ Dong Yang ⋅ Zhengyi Lu ⋅ Junchao Zhu ⋅ Tianyuan Yao ⋅ Juming Xiong ⋅ Junlin Guo ⋅ Yanfan Zhu ⋅ Yuechen Yang ⋅ Daguang Xu ⋅ Bennett Landman ⋅ Yucheng Tang ⋅ Yuankai Huo
Abstract
Vision-language models (VLMs) applied to 3D medical scans face a structural memory wall. A single CT or MRI volume produces tens of thousands of vision tokens, and the resulting key-value (KV) cache scales linearly with both batch and context. Existing KV compression methods operate on the cache's feature or token statistics, or apply 1D and 2D transforms, leaving the three-axis spatial structure in the cache unexploited. We propose TACT-KV (Tri-Axis Cosine Transform), a calibration-free KV compression method that exploits strong spatial correlations in visual KV states along all three anatomical axes of 3D scans. This correlation concentrates the KV cache energy under a frequency-domain transform. TACT-KV applies a separable 3D discrete cosine transform to the KV cache, allocates a per-layer bit budget across frequency bins via reverse water-filling (more bits to high-energy components), and quantizes each bin with a fixed scalar codebook. We derive a close-form relation: the distortion reduction from adaptive bit allocation is dominated by the spectral concentration of the transformed cache. Across six medical VLMs and two 3D CT benchmarks, TACT-KV maintains near-lossless prediction fidelity at 1-bit. The same advantage reproduces on four general-purpose VLMs evaluated on video tasks, indicating that the benefit of exploiting three-axis structure generalizes beyond medical imaging. At equal high-bandwidth memory (HBM), TACT-KV scales the decoding batch size on a single GPU by up to $8\times$.
Chat is not available.
Successful Page Load