Can Cross-Model KV Cache Mappings Survive Longer Prompts? An Empirical Analysis
Rishi Khare ⋅ Jiayi Qian ⋅ Zishen Wan
Abstract
Reusing a KV cache across models would let a small, cheap model perform prefill while a large model decodes, cutting the dominant cost of serving long prompts. Recent work shows this is possible within a model family, via a closed-form linear map between models of different sizes \citep{heo2026crossmodel}, so we test whether such a map preserves accuracy in long-context settings. Evaluating all four directed transfers across two Qwen3 pairs, we find that the technique's value inverts between short and long context settings. On short-context multiple choice benchmarks (ARC-C and HellaSwag), a small-to-large mapped cache consistently outperforms serving the source model alone, by up to $8.40$ points. However, on all eight long-context retrieval configurations (RULER SQuAD and HotPotQA at $4096$ and $8192$) small-to-large cache transfer falls below the smaller model baseline in task quality, by up to $14.76$ points, which indicates that a deployment would provide better answers by ignoring the large model entirely. We further find that large-to-small transfer, which prior work leaves largely unmeasured, retains at least as much of its target in all configurations and materially more on long-context settings compared to small-to-large transfers.
Chat is not available.
Successful Page Load