Which Features Matter? Memory-Efficient Compression for Cross-Model KV Mappings
Abstract
Recent work has shown that cross-model KV cache transfer can be performed via closed-form affine maps within the same model family. By fitting an affine map from a source model's KV cache to a target model's cache using ridge regression, recent work has demonstrated that cross-model transfer avoids redundant prefill in multi-agent LLM serving, though these maps can consume 4-12GB per model pair [Heo et al., 2026]. We ask whether the KV cache affine mapping itself can be compressed while preserving the downstream task quality of the target model. Applying reduced-rank regression, we find that key maps have more aggressive rank compression (within 3pp of the dense map at half rank) than value maps, which collapse in accuracy (up to over 21pp) at the same rank. Further, we show that compressing the map by calibrating on its fitted responses to input data retains substantially more accuracy than truncating the weight matrix by SVD because SVD optimizes for the weight matrix's own singular directions rather than the directions that explain its fit on the calibration data. Through these results, we identify the asymmetry between applying rank compression on key and value maps and show the cross-model KV map's memory consumption can be reduced while retaining task quality via calibration-aware truncation.