Perplexity Cost Understates What Activation Quantisation Breaks
Abstract
Efficient inference is what makes deploying large language models feasible outside well-resourced data centers: post-training quantisation is the standard way to shrink a model's memory and compute footprint before it is served on constrained hardware. Whether a quantised model still works is almost always answered with a single number, perplexity, and a small perplexity cost is read as evidence the model is safe to deploy. We show this can be badly misleading. Treating the residual-stream activation a quantiser transforms as a structured object rather than a scalar feeding a loss, we pair three in-context probes (copying, copying across a span, and retrieval of a value bound to a key) with a matched noise control across 12 models from four families. Perplexity ranks quantisation methods reliably, but where it has risen by only a factor of 1.2 to 1.5, a deployed model can still have lost half its retrieval accuracy while its copying ability is untouched, a gap a perplexity-only check would never surface. This is not an artefact of a diagnostic method: the same split reappears under three quantisers actually used in practice today (GPTQ, AWQ, SmoothQuant) once they compress activations, and it holds up to 32B parameters. We trace the cause to a specific, correctable pattern in how the quantisation error is structured, and show that a deployable fix, rotating the representation before quantising with QuaRot's construction, restores the lost capability at a measured latency cost. For teams deploying compressed models where compute is scarce and the ability to catch failures after the fact may be limited, the lesson is direct: a perplexity target is not sufficient evidence that a cheaper model still does what it is trusted to do, and cheap, targeted behavioural checks can catch failures perplexity misses before deployment. Code is available at https://anonymous.4open.science/r/Perplexity-Cost-Understates-What-Activation-Quantisation-Breaks-8712/README.md