A Calibrated Canary: Per-Input Verification for Sub-8-Bit KV Compression in On-Device Agents
Abstract
An on-device agent’s standing memory cost is its key–value (KV) cache, and compressing it below 8 bits is free on one model and catastrophic on another. Offline, per-model schemes cannot see that, with weights fixed, the safe policy depends on the input: tool-call traces are the safest input for two Llama models and the most damaging for Qwen2.5-3B. We instead run one probe forward that applies the candidate compression to the live input, then calibrate it: damage tracks the reading by a measured ratio κ, so an operator declares a quality budget q and the threshold κq follows. Across 75 paired cells the gate holds 74/75 inside a 1.2% budget where a fixed-threshold fuse holds 57/75, at a mean compression cost of 2.01× → 1.94×. A probe that stops once the threshold’s side is settled costs 54% of a full probe and matches its decision on 74/75 cells, with no permissive disagreement, on corpora its calibration never saw. Two results are negative: a graded ladder atop the calibrated threshold adds at most 0.02× compression and costs two cells of adherence at the tightest budget, and the gate declines 1.72× of compression it had already measured as safe.