The Prefill Ceiling: Why Quantizing the KV Cache Costs On-Device Agents Context
Abstract
On-device agents built on small language models read new material when they retrieve documents, tools, or files. Every read is an uncached prefill, so the prefix cache that normally sustains decode can not be used. We show that the standard solution of quantizing the KV cache actually shortens the context an Apple Silicon device can process as its fused attention kernel refuses quantized caches. Thus, prefill falls back to materializing the complete score matrix at query-head count. On mlx-lm this costs 256 KiB per context token, which outweighs the 92 KiB cache saving from quantization. We determined that since the score matrix grows with the number of query heads and the cache savings grow with the number of KV heads, the winning effect is determined by the model's architecture. We state this as a rule, and it predicted peak memory to about 1% across contexts of 8k to 65k tokens. It also accurately predicted a multi-head model on mlx will save memory instead. A tiled prefill that removes the score matrix processes 3.5x the context of fp16 and 7.7x of the stock quantization. Of the nine real documents swept, this research determines that 5 are processable only with the tiled prefill we implement. A 28k token paper runs at fp16, but is over the device's wired limit once quantized. With three concurrent 3B agents, non-tiled (stock) quantization causes 1.98 GB of memory compression, versus the 0.27 GB of tiling.