SMASH: Probing Speech Recognition Robustness via Semantically Targeted Bit Flips
Zafaryab Haider ⋅ Md Hafizur Rahman ⋅ Aysegul Bumin ⋅ Prabuddha Chakraborty
Abstract
Automatic speech recognition (ASR) systems are usually judged by aggregate transcript accuracy, but real failures often hinge on meaning: for example, a wrong dose, address, account number, or date can matter more than many harmless word errors. We introduce SMASH, a framework for finding targeted semantic failures caused by sparse INT8 weight-bit perturbations, a dominant model data type for ASR edge devices, in quantized sequence-to-sequence ASR models. Given an audio input and a meaning-critical span,SMASH searches for small weight changes that replace the intended meaning while keeping the transcript readable, plausible, and close to the original output. Each accepted fault is a certificate that a specific semantic substitution is reachable under a bounded perturbation budget. Empirical study focuses on numeric and scalar substitutions across three ASR backbones: Whisper-small.en (WS), Whisper-large-v3 (WLv3), and SeamlessM4T-v2-large (S-M4T); and three corpora: an LLM-assisted controlled corpus (LACC), MultiMed, and LibriSpeech. On LACC at a 20-bit budget ($B$), SMASH-hybrid accepts numeric substitutions on 24/45 WS targets, 19/44 WLv3 targets, and 5/34 S-M4T targets; reachability is lower on real-speech datasets. In matched LACC budget sweeps tested up to $B=40$, WS reaches its ceiling already at $B=20$, while WLv3 and S-M4T continue to gain accepted targeted substitutions through $B=40$. Human validation supports the gate-based labels: on a 105-item shared validation set, independent human-majority labels and a four-judge large- language-model ensemble agree on the composite clean_success label, with $\kappa=1.00$. The faults are sparse: the pooled median is three INT8 bit flips, with medians of two flips on WS, eleven on WLv3, and twelve on S-M4T. Entity and negation pivots are evaluated as secondary evidence in the appendix. SMASH shifts ASR robustness evaluation from "how much did the transcript change?'' to "which meanings can be changed, how plausibly, and with how few bit flips?''
Chat is not available.
Successful Page Load