When Speech Errors Break Tasks: Cross-Layer Evaluation of Conversational Voice Agents
Abstract
Small-model agent stacks delegate recurring, well-scoped operations to compact specialist modules. In a voice agent, speech recognition is one such module: its transcript becomes the input to routing, typed argument handling, and dialogue state. Selecting a recognizer by Word Error Rate (WER) alone can therefore miss whether the module preserves the downstream task contract. We introduce a cross-layer evaluation protocol centered on two complementary diagnostics. Pragmatic Intent Failure (IntentFail) records whether a version-pinned router assigns different routes to a reference transcript and its ASR hypothesis, while Slot & Entity Error Rate (SEER) tests whether typed reference values remain recoverable after spoken-form normalization. Our industry-grounded evaluation suite distills handler relations, intent cues, dialogue patterns, and entity formats from access-controlled real-world interactions, instantiates new evaluation records under this blueprint, and applies task-specific quality gates. We compare six compact ASR systems with nominal sizes from 0.6B to 3B on 336 speech-rendered customer-service turns. Under degraded audio, raw WER, normalized WER, and IntentFail favor different point-estimate leaders, although routing confidence intervals overlap. Pooled over 3,996 scored pairs, 49% of fixed-router changes occur at zero normalized WER. On 400 entity-rich turns, tracking-number SEER ranges from 18.8% to 34.4% and order-ID SEER from 7.2% to 44.6%, compared with 0.4% to 3.0% for dates. The results provide a task-integrity axis for comparing compact specialists within heterogeneous agent stacks.