The Price of Precision: A Calibrated Cost Model for Mixed-Precision SLM Serving
Abstract
Small language models (SLMs) are increasingly served on-device, where quantization is not optional but the default, and where an agentic session sweeps its operating point as tool-call context grows. Choosing a mixed-precision configuration trades speed against accuracy, and prior allocators price that speed by a proxy such as the average bit-width or a memory budget. Served speed, however, is governed by the binding resource (compute, weight memory, or the key-value (KV) cache), which shifts with phase, batch, and context length, so fewer bits can be slower. We build a single cost model of served mixed-precision speed that reasons about the binding resource directly, and put it to two uses. First, as a predictor: given a configuration and operating point, it forecasts absolute served speedup. Calibrated to a consumer Radeon 780M accelerated processing unit (APU) serving integer GGUF quants, it predicts held-out decode within 9.4% on a dense Llama-3.2-1B and 7.1% on the mixture-of-experts (MoE) OLMoE-1B-7B, and it captures a sign inversion (fewer bits are sometimes slower) that the bit-count proxy get wrong. Second, we invert the same model into a per-tensor allocator: we model its per-unit accuracy and time costs as additive, and use an exact dynamic program to return the optimal configuration at a target served-speed budget or accuracy budget. Turning the decision form into absolute time adds a few constants calibrated once per device. Because an agent's KV cache grows turn by turn, the binding resource shifts within a single deployment, and the optimal configuration provably shifts with it.