When Speech Errors Break Tasks: Cross-Layer Evaluation of Conversational Voice Agents
Abstract
Voice agents increasingly convert speech into state updates and actions under open-ended language and variable acoustic conditions. A transcript can remain lexically close to its reference while changing the selected route or a structured argument, creating a silent interface failure that component metrics do not localize. We introduce a cross-layer evaluation protocol that holds downstream consumers fixed and traces speech errors from transcript edits to route changes, typed-value loss, and dialogue-state outcomes. IntentFail measures changes under a fixed intent router, while SEER measures recovery of typed reference entities. We build an industry-grounded suite through a governed three-stage process: a model distills handler relations, intent cues, dialogue patterns, and entity formats from access- controlled real-world interactions; separate prompts instantiate new records with fictional values; and quality gates enforce label clarity, spoken style, required fields, and trajectory consistency. Across six ASR systems, 336 speech-rendered turns, and a 400-turn entity track, lexical and task-interface diagnostics expose different system behavior. Pooled across scored systems and acoustic conditions, 49% of fixed-router changes occur at zero normalized WER. Tracking-number SEER ranges from 18.8% to 34.4% and order-ID SEER from 7.2% to 44.6%, compared with 0.4% to 3.0% for dates. Experiments with 462 composed dialogues and three synthesis systems extend the same audit across state and generated-speech boundaries. These cross-layer signals make failures at the speech-to-action interface observable and motivate targeted verification before an agent commits state or passes arguments downstream.