Compression has an accent: aggregate PTQ benchmarks hide who pays for on-device speech recognition
Bruce Peters ⋅ Diego Mardian
Abstract
Post-training quantization (PTQ) is one of several ways to fit a speech foundation model on a device, alongside quantization-aware training and distillation, and the cheapest, because it needs no retraining. The speech PTQ benchmarks we found check accuracy with an aggregate error rate, and we ask whether that is a safe reliability check. Six non-native accent groups and native speakers read the same $150$ sentences, and we measure PTQ per speaker group for three encoders across two architecture families (wav2vec2, Zipformer) and two quantizers (static INT8, 4-bit weight-only). About $96\%$ of the native--Vietnamese difference already exists before compression ($84\%$ on a common noise floor). The part compression adds grows with the accuracy a recipe spends and lands on speakers the model already finds hard ($r{=}.95$), although random weight noise reproduces both relations, so neither is specific to quantization. The ordering, the marginality relation and the INT8 amplification hold again on a second corpus with $18$ to $373$ speakers per group. On one encoder, the 4-bit model reads as lossless only because native speech appears to improve, and $70\%$ of that improvement is the model no longer transcribing flaps that native speakers produce and the canonical reference cannot encode. The aggregate is lossless because of how the compression interacts with the reference, which a per-group breakdown also misses. Calibrating on a group does not detectably protect it ($n{=}4$ per group), and certifying static PTQ on one outlier-sensitive calibration draw is unsafe, since three MinMax draws from one pool gave $\Delta$PER $.380$, $.379$ and $.908$. On an iPhone 16 the English wav2vec2 loads only when quantized and then runs in under half real time, and its INT8 outputs, unlike its 4-bit ones, change with the runtime build.
Chat is not available.
Successful Page Load