MUTE: Multi-Level Alignment Uncoupling Against Talking-Head Exploitation for Voice Protection
Abstract
Talking-head generation models, which synthesize realistic facial animations from audio, are increasingly vulnerable to misuse in multimodal deepfake scenarios. Protecting such systems remains challenging, as talking-head models fundamentally rely on precise audio–visual alignment, while effective audio protection methods remain largely underexplored. In this work, we propose MUTE, an audio protection framework that uncouples the underlying audio–visual alignment through a multi-level strategy. Specifically, MUTE combines (1) Representation-level Degradation, which perturbs temporal audio embeddings to degrade their representations, and (2) Alignment-level Disruption, which directly perturbs cross-modal attention to disrupt structural alignment. To improve robustness, perturbations are constrained in the STFT domain to high-energy regions, making them resistant to post-processing and denoising. An optional speaker-level objective further mitigates potential bypass via TTS-based resynthesis. Extensive experiments demonstrate that MUTE consistently degrades lip synchronization across both white-box and black-box talking-head models while preserving perceptual audio quality. The proposed method remains effective under real-world transformations and can be combined with image-based approaches to provide stronger multimodal protection.