When Speech Errors Break Tasks: Cross-Layer Evaluation of Conversational Voice Agents
Abstract
Real-time conversational agents can respond fluently while speech processing silently changes the route or argument passed downstream. WER and MOS characterize component quality but do not directly measure whether task-critical content survives across system boundaries. We introduce task-content integrity, a cross-layer evaluation framework centered on two diagnostics: IntentFail measures changes under a fixed intent router, and SEER measures recovery of typed reference entities. We construct an industry-grounded evaluation suite by distilling task, dialogue, and entity structure from governed real-world customer-service interactions, then using separate prompts to generate new evaluation turns with fictional values. Across six ASR systems, 336 TTS-rendered turns, and a 400-turn entity-rich track, WER, normalized WER, and IntentFail select different point-estimate leaders under controlled acoustic degradation. Pooled across systems and conditions, 49% of fixed-router changes occur at zero normalized WER. Tracking-number SEER ranges from 18.8% to 34.4%, and order-ID SEER from 7.2% to 44.6%, compared with 0.4% to 3.0% for dates. Experiments with 462 composed dialogues and three TTS systems extend the same evaluation contract across dialogue-state and synthesis boundaries. These results show that cross-layer task-content integrity provides a practical complement to lexical, perceptual, timing, and interaction metrics for real-time multimodal conversational AI.