The Prefill Ceiling: Quantizing the KV Cache Shortens Usable Context on Apple Silicon
Abstract
On-device inference is motivated by a variety of factors. Reading a contract or code base locally for privacy, are such examples. However, on-device inference is severely bounded by memory. The standard solution of quantizing the KV cache actually shortens the context an Apple Silicon device can process as its fused attention kernel refuses quantized caches. Thus, prefill falls back to materializing the complete score matrix at query-head count. On mlx-lm this costs 256 KiB per context token, which outweighs the 92 KiB cache saving from quantization. We determined that since the score matrix grows with the number of query heads and the cache savings grow with the number of KV heads, the winning effect is determined by the model's architecture. We state this as a rule, and it predicted peak memory to about 1% across contexts of 8k to 65k tokens. It also accurately predicted a multi-head model on mlx will save memory instead. A tiled prefill that removes the score matrix processes 3.5x the context of fp16 and 7.7x of the stock quantization. Of the nine real documents swept, this research determines that 5 are processable only with the tiled prefill we implement. A 28k token paper runs at fp16, but is over the device's wired limit once quantized.