KVPort: Zero-Prefill Cross-Family KV Cache Transfer Across Tokenizers Without Receiver Prefill
Abstract
Routing requests between language models forces the receiver model to rebuild its key-value (KV) cache requiring a full prefill. Heo et al. demonstrated that a linear mapping can be used to transfer KV caches to other language models belonging to the same model family. We extend the cache transfer across families, where the tokenizers are different by introducing KVPort. KVPort achieves this by aligning tokens by their UTF-8 byte spans, converting rotary position embeddings from the source's frame to match the receiver's frame, and then mapping each key and value between models using a per-head ridge regression. Across Qwen3-8B and Llama-3.1-8B, KVPort enables effective cross-family KV cache transfer while skipping receiver prefill.