Perfect transcription, opposite verdicts: ASR-derived rewards carry verifier-specific prosody preferences
Abstract
Conversational voice systems increasingly use automatic speech recognition (ASR) to check whether generated speech says the intended words. The same verifier may rank candidates, filter training data, or supply a reward, thereby affecting not only intelligibility but also which deliveries survive. We audit this measurement assumption on 19,567 expressive–neutral pairs from 125 English speakers with nine ASR verifiers from six training lineages. Word error rate (WER) worsens for expressive delivery in all 27 verifier–corpus cells, but reference-transcript negative log-likelihood (NLL) gives opposite reward directions: two verifiers reward expressive delivery while seven penalise it. Their mean penalty rates are 39.1% and 63.6%, a 24.5-point gap; filtering to pairs transcribed perfectly on both sides leaves 22.6 points. The result survives loudness and duration controls, while effect sizes qualify it: the strongest non-overlapping reversals occur in the two corpora with least text variation, and encoder–decoders are near null on ESD-en. Frame normalisation changes the direction of three of seven alignment-based systems, so architecture is not established as causal. A repeated-take null control returns chance, and a causal perturbation shows that expanded energy variance moves two CTC verifiers in opposite directions. Accuracy, scale, and the measured acoustic correlates do not explain the panel, and pooling is not a demonstrated remedy. These are offline experiments on human recordings. We do not evaluate synthetic candidates, streaming latency, turn-taking, multimodal fusion, a live agent, or a downstream selection or training loop. The result is therefore an evaluation-foundations warning for expressive conversational AI, not a demonstrated system-level harm.