Quantizing the Judge: Verdict Drift Under LLM.int8 and NF4
Abstract
Quantizing an LLM judge is normally treated as a systems optimization, yet it may change the measurement instrument itself. We compare two fixed 7B judges (Qwen2.5-7B-Instruct and Vicuna-7B-v1.5) under fp16, BitsAndBytes LLM.int8, and NF4 on human-labelled Chatbot Arena pairs, scoring every pair in both orders. Across three registered strata, INT8 flips 8.5--11.1% of fp16 verdicts while its human-agreement changes remain within a pre-registered +/-5-point equivalence region. NF4 flips 23.8--31.3%, changes position consistency by -11.2 points for Qwen and +7.5 to +14.0 for Vicuna, and lowers Vicuna agreement by 10.7 points (95% CI: -12.7, -8.8) on a Vicuna-labelled identity subset. Identity-affinity changes are unresolved. NF4 reduces peak allocated memory by about 58% with fp16-like latency; INT8 reduces it by 38--40% but is slower in this environment. One NF4 stratum has 98.8% parse coverage, below our 99% certification threshold, and is explicitly qualified. Thus the tested deployment recipes are not behaviorally interchangeable, even when aggregate agreement appears stable.