Right Tool, Wrong Arguments: Permission Drift in Quantized Tool-Calling LLMs
Abstract
Post-training quantization is the standard way to deploy tool-using language model agents on constrained hardware, and is generally assumed safe once it preserves benchmark accuracy. We test that assumption on four safety-aligned, tool-calling model families (Qwen, Llama, Mistral, Granite) at four precisions (BF16, FP8, INT8, INT4). On the Berkeley Function-Calling Leaderboard, three of the four families lose at most 2.17 accuracy points at INT4 and one improves by over 11 points: ordinary tool-calling capability survives quantization. A paired, case-level analysis of the same predictions tells a different story — up to 13.50\% of individual decisions change identity between BF16 and INT4, even when the aggregate score is stable or improves. We then introduce a permission-boundary analysis that gives each model an explicitly authorized tool call and measures whether an embedded, plausible unauthorized instruction changes the arguments it supplies. Every model family exhibits at least one reproducible case where a securely authorized action at BF16 becomes unauthorized at INT4, and INT4 is consistently the least behaviorally stable precision, changing 20--50\% of records relative to BF16. A seven-category breakdown of these violations — spanning recipients, permissions, credentials, and financial and destructive operations — shows the effect is concentrated unevenly and inconsistently across models, with no single category or precision behaving predictably from one model to the next. We study this failure mode, called \textit{permission drift}: quantization can leave a tool-using agent's benchmark accuracy intact while silently moving the arguments it is willing to submit outside the user's authorization boundary. Aggregate accuracy is therefore not a sufficient signal that a quantized agent is safe to deploy with real permissions.