What Was That Again? Certified Robustness for Automatic Speech Recognition
Abstract
Automatic Speech Recognition systems are notoriously both sensitive to input perturbations and challenging to defend, which is a product of their high-dimensional, discrete output spaces. Traditional Randomized Smoothing workflows collapse in sequence-to-sequence tasks because the probability mass of any single transcription vanishes under noise. We propose an Anytime-Valid Certified Transcription framework that replaces fixed-sample binomial testing with E-value Martingales. Our approach leverages a dual-gate pipeline: a Two-Sided Atomic Audit that accumulates statistical wealth to certify both token existence and adversarial exclusion, and a Rank-Based Tournament that selects the winning sequence. Our evaluations across four diverse architectures demonstrate that this approach yields an average 25.5\% relative reduction in Word Error Rate, while also providing granular word- and sentence-level certifications to enhance acoustic security.