Can Muted Talking Faces Guide Expressive TTS? A Controlled Study of Visual and Lexical Cues
Abstract
Can visible performance guide expressive speech generation without using source speech as an expressive prompt? We test a compact interface between frozen models: Gemini converts text, a muted talking-face clip, or both into a structured delivery instruction, and VoxCPM2 renders that instruction. A deterministic corpus search yields 169 English MAFW clips with exactly one acoustic speaker and one visible active speaker. With one independent voice reference, visual+text reduces matched-original normalized DeEAR error relative to visual-only by 0.0308 (95% bootstrap CI [0.0064, 0.0550], Holm p=.015), but not relative to text-only (0.0099, CI [-0.0148, 0.0345], p=.381). A second 507-waveform run with per-clip speaker controls reproduces the aggregate visual+text–visual advantage, although an audit shows that these controls retain source-ranked prosody and arousal. Thus, lexical content reliably disambiguates muted visual behavior, but the automatic evidence does not establish an incremental visual benefit beyond text. We report the positive contrast, negative boundary, and control leakage together as a controlled result for language-mediated multimodal speech generation.