Can the Query Still Find It? Argmax Recoverability as a Label-Free Screen for KV-Cache Compression
Shrenik Bhansali ⋅ Kumar Gaurav ⋅ Amol Ambardekar ⋅ Larry Heck ⋅ Karthik Vijayan
Abstract
Deploying a long-context language model requires choosing a key--value (KV) cache compression configuration under a memory budget, yet configurations of nearly identical size can have opposite downstream outcomes. We distinguish two principal failure modes: compressed keys can send a query to a different token than the full-precision cache would, or compressed values can corrupt a token that is still selected correctly. We study the first with \emph{argmax recoverability} $\AR@b$, the fraction of unlabeled probes on which the token ranked first by a full-precision query remains among the top-$b$ tokens scored with compressed keys. It requires no generation and no reference answers. It also separates mechanisms: pure token deletion produces a flat recoverability curve, whereas bounded-distortion codes recover demoted routes as $b$ grows. Across $45$ matched-memory configurations on two architectures, $\AR@{16}$ selects the better member of a near-equal-memory pair $93\%$ of the time, against $57\%$ for the memory budget. That gap has structure. Within a single compression family the memory budget is a reasonable guide ($86\%$); across families it is no better than chance ($49\%$). $\AR@{16}$ is correct $95\%$ and $93\%$ of the time respectively. A threshold calibrated at 4K transfers without recalibration to a held-out architecture and, for retrieval, to contexts up to 64K. $\AR@b$ is therefore useful as a first-stage screen rather than a replacement for benchmarking.
Chat is not available.
Successful Page Load