When Speech Errors Break Tasks: Cross-Layer Evaluation of Conversational Voice Agents
Abstract
Robustness evaluation for interactive AI requires failure descriptors that connect low-level perturbations to downstream behavior. In voice agents, aggregate transcript error can obscure whether a speech difference changes a task route or corrupts a typed value. We introduce a cross-layer evaluation protocol with two complementary descriptors. Pragmatic Intent Failure (IntentFail) records whether a version-pinned router changes its route between a reference transcript and an ASR hypothesis. Slot & Entity Error Rate (SEER) tests whether typed reference values remain recoverable after spoken-form normalization. An industry-grounded construction process distills handler relations, intent cues, dialogue patterns, and entity formats from governed real-world interactions, then uses separate prompts and quality gates to create new evaluation records with fictional values. Across six ASR systems on 336 speech-rendered turns, raw WER, normalized WER, and IntentFail favor different point-estimate leaders under degraded audio, although routing confidence intervals overlap. Pooled over 3,996 scored pairs, 49% of fixed-router changes occur at zero normalized WER. On 400 entity-rich turns, failure rates vary sharply by value type. These cross-layer descriptors and error patterns provide an evaluation substrate that can support future robustness search across speech-agent interfaces.