Compression Breaks Discipline: Quantization Increases Missed Tool Calls in Small Language Models, but Not Mid-Size Ones
Muhammad Roshaan ⋅ Muhammad Aaliyan ⋅ Saifullah Bin Zubair
Abstract
Small language models are increasingly deployed as tool-calling agents under tight memory budgets, where 4-bit post-training quantization (NF4) is the standard compression choice. Prior evaluations of quantized models report aggregate accuracy, which can hide which kinds of failures change even when the overall score appears stable. We audit the error distribution of tool-calling behavior under quantization using a seven-class error taxonomy over 600 controlled generations from two Qwen3 models (1.7B and 4B), each tested at FP16 and NF4 on Berkeley Function Calling Leaderboard single-turn tasks with identical greedy decoding. Our central finding is a size-dependent behavioral shift: for Qwen3-1.7B, NF4 nearly triples missed required tool calls (6.0% to 17.3%; $\Delta=-11.3$pp, paired 95% CI $[-17.3,-6.0]$, McNemar $p=0.0001$), spread across all task categories; for Qwen3-4B the same comparison shows no effect (7.3% vs. 7.3%, paired CI $[-4.7,+4.7]$) with nominally higher accuracy. Aggregate accuracy moves by only a few points and, in the 4B, points the wrong way. A direct paired interaction test confirms the asymmetry ($-11.3$pp, 95% CI $[-20.0,-3.3]$, $p=0.006$), and 13 of the 26 NF4 misses were answered correctly under FP16. These results suggest that quantization damage in this family manifests as a loss of tool-calling discipline rather than knowledge, and that 4-bit safety should be validated per checkpoint instead of assumed from a size threshold. We release our taxonomy classifier, generation protocol, per-answer labels, and paired-analysis code to support error-profile audits beyond scalar accuracy.
Chat is not available.
Successful Page Load