A Multilingual Audit of Numeric Retrieval Failures Under KV-Cache Eviction
Abstract
Retrieval accuracy is the standard summary of what KV-cache eviction costs, but it does not say how a policy fails when it does. We audit three eviction policies (KNorm, SnapKV, and StreamingLLM) on two 7–8B models across seven languages, using OneRuler's present- and absent-needle numeric retrieval tasks, and separate whether an output contains a numeric candidate from whether that candidate is correct. Policies with nearly identical accuracy can fail in opposite ways: at 25% retention on Llama under query-agnostic compression, KNorm emits a candidate on most present trials but is usually wrong, whereas StreamingLLM emits far less often and is almost always correct when it does, yet still emits on a third of absent trials. KNorm's errors have a distinctive signature: nearly all of its wrong numeric outputs contain a digit truncation of the true value. A query-aware protocol restores SnapKV to near-baseline accuracy and reorders the policies, although the KNorm–StreamingLLM contrast persists. Finally, two properties of the released benchmark change multilingual conclusions. First, the Korean haystack repeats a short source text about four times; Korean, Llama's weakest language uncompressed, ranks among the strongest under the content-based policies but not under position-based StreamingLLM, and excluding it removes the evidence for a headline cross-language dispersion effect. Second, the absent-task scorer credits outputs that list distractor values before declaring absence, raising Llama's Korean absent-task score from 0.12 to 0.83. We recommend reporting emission and conditional correctness alongside accuracy, together with the serving protocol, corpus-repetition checks, and an explicit output contract.