Model Selection and KV-Cache Allocation Are Coupled under Memory Constraints
Abstract
For on-device long-context inference, model selection determines the memory available for the KV cache, and cache allocation changes model quality under a fixed device-memory budget. We study this interaction across 15 models from seven architecture families, at context lengths from 4K to 128K tokens. Increasing the context length repeatedly changes the highest-scoring model under a fixed budget. These changes arise because weight memory does not predict KV bytes per token, so models require substantially different eviction fractions under the same budget. Uncompressed ranking orders models by their uncompressed score and then evicts cache from the chosen model until it fits. On 100 additional GovReport documents, this policy loses 4.02–9.44 points relative to budget-aware ranking, which scores every model at the eviction fraction its budget requires. At long context, the required eviction can reverse a ranking that holds at shorter contexts. Model selection and cache allocation should be decided jointly for each context length.