When Speech Errors Break Tasks: Cross-Layer Evaluation of Conversational Voice Agents
Abstract
Task-oriented voice agents map free-form speech into an operational task representation: a routed category and typed reference values. Surface transcript metrics do not directly test whether this representation remains stable after a speech boundary. We introduce a cross-layer evaluation protocol with two complementary diagnostics. Pragmatic Intent Failure (IntentFail) records when an ASR hypothesis changes a version-pinned router's output, while Slot & Entity Error Rate (SEER) tests whether typed reference values remain recoverable after spoken-form normalization. An industry-grounded construction process distills handler relations, intent cues, dialogue patterns, and entity formats from governed real-world interactions; separate prompts instantiate new records with fictional values, and quality gates screen clarity, ambiguity, spoken style, required fields, and trajectory consistency. Across six ASR systems on 336 TTS-rendered customer-service turns, raw WER, normalized WER, and IntentFail produce different nominal leaders, although routing confidence intervals overlap. Pooled over 3,996 scored pairs, 49% of fixed-router changes occur at zero normalized WER, showing that normalization equivalence does not guarantee invariance at the task interface. On 400 entity-rich turns, tracking-number SEER ranges from 18.8% to 34.4% and order-ID SEER from 7.2% to 44.6%, compared with 0.4% to 3.0% for dates. Manual inspection of 60 target-gated route changes identifies recurring critical-word substitutions and sensitivity to orthographic and punctuation form. Composed-dialogue and synthesis-boundary probes extend the same structured contract across state and generated speech. The results motivate evaluation that separates surface fidelity, task-category stability, and typed-value integrity.