Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation
Abstract
Recent progress in speech-driven 3D facial animation has improved vertex level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained lip jaw coordination and the mapping from acoustics to motion is inherently one to many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional lip motions and composes them into geometry consistent facial deformation. To represent visible articulation with three canonical directional motions, spreading, opening, and protrusion, we propose a Speech Articulatory Memory (SAM) that captures the correspondence between speech and directional articulatory motions under phonetic context through key-value memory structure based retrieval and decoding. Then, a Topology-aware Articulatory Composition (TAC) integrates the predicted directional motions under mesh topology to produce coherent 3D facial motion. Experiments on VOCASET and TFHP show that our method achieves the-state-of-the-art performance on standard reconstruction metrics and improves visible articulatory distance and velocity fidelity for key lip factors, while a user study confirms clear preference in lip sync and realism.