Right Tool, Wrong Arguments: Permission Drift in Quantized Tool-Calling LLMs
Abstract
Post-training quantization is the standard way to deploy tool-using language model agents on constrained hardware, and is generally assumed safe once it preserves benchmark accuracy. We test that assumption across four tool-calling model families (Qwen, Llama, Mistral, and Granite) at four precisions (BF16, FP8, INT8, and INT4). On the Berkeley Function-Calling Leaderboard, three families lose at most 2.17 accuracy points at INT4 and one improves by over 11 points: aggregate tool-calling capability largely survives quantization. A paired, case-level analysis of the same predictions tells a different story---up to 13.50\% of individual decisions change between BF16 and INT4, even when aggregate accuracy remains stable or improves. We then introduce a permission-boundary analysis that gives each model an explicitly authorized tool call and measures whether untrusted external content changes the arguments it supplies. Every family exhibits at least one reproducible case where a securely authorized BF16 action becomes unauthorized at INT4, and INT4 changes 20--50\% of permission records relative to BF16. Argument churn also exceeds tool-selection churn across all four families. These transitions are deterministic under repeated greedy decoding, although none of the 36 primary security comparisons survive Holm correction; the evidence therefore establishes case-level instability rather than population-wide average degradation. We call this failure mode \textit{permission drift}: quantization can leave aggregate tool-calling accuracy intact while silently moving sensitive arguments outside the user's authorization boundary.