Emotion-Aware Talking Face Generation via Audio-Visual Feature Aggregation
Abstract
Audio-driven talking-face generation has achieved increasingly accurate lip synchronization; however, generating natural facial expressions while preserving the identity of unseen speakers remains challenging. Existing approaches often focus primarily on speech-related mouth movements, while emotional facial expressions are either neglected or require additional identity-specific training or reference videos [1]. We propose an emotion-aware face-talking generation framework that generates expressive, lip-synchronized videos of arbitrary identities from an audio signal, a single reference image, and an emotion specified in text form. Our approach models facial motion as the interaction of four factors: facial geometry, speech content, speaker identity, and emotion. Facial landmarks extracted from the reference image are represented as a graph and encoded using a transformer-based graph neural network with hierarchical pooling to preserve facial geometry and capture relationships between facial regions. An audio-disentanglement network separates speech-content and speaker-identity information. Wav2Vec is employed to obtain speech representations from which identity- and emotion-independent content features are extracted [2]. Speaker identity is jointly learned from audio and the reference image by projecting their representations into a shared latent space. The desired emotion is provided as text and encoded using a pretrained transformer. Landmark, speech-content, identity, and emotion representations are aggregated to predict facial landmark displacements. A temporal attention mechanism further considers sequences of predicted landmarks to reduce noise and sudden inter-frame movements. Finally, an image-generation and refinement module transforms the predicted landmarks and reference image into high-quality video frames while preserving identity and emotional characteristics. The model was evaluated on the MEAD and CREMA-D emotional audio-visual datasets and compared with ATVG, MakeItTalk, Audio2Head, and EAMM. On MEAD, our approach achieved an SSIM of 0.708, PSNR of 30.621, CPBD of 0.058, M-LMD of 3.221, and LMD of 3.282, outperforming the comparison methods across all reported metrics. Similar improvements were observed on CREMA-D, where the proposed model achieved an SSIM of 0.712, PSNR of 32.032, CPBD of 0.066, M-LMD of 4.512, and LMD of 3.956. Ablation experiments further demonstrated the contributions of the graph-based landmark representation, audio disentanglement, emotion encoding, temporal attention, and image-refinement components. A user study with 20 participants additionally evaluated video quality, lip synchronization, identity preservation, and emotion representation. The proposed approach obtained mean opinion scores of 4.01, 4.29, 4.45, and 4.51, respectively, outperforming the evaluated state-of-the-art approaches in all four categories. These results demonstrate that explicitly aggregating geometry, speech, identity, and emotion representations can generate expressive talking-face videos while maintaining accurate lip synchronization and the characteristics of previously unseen identities.