From Spectrograms to Shallow Speech SSL: Audio-Frontend Design for Generator-Disjoint and Low-Latency Audio--Visual Deepfake Detection
Qiwu Wen ⋅ Md Jahangir Alam
Abstract
Video-mediated conversational systems increasingly rely on synchronized speech and facial video, making multimodal input integrity important for reliable perception. Audio--visual deepfake research has mainly improved cross-modal fusion, whereas the audio representation supplied to the detector remains comparatively underexamined. We study this design choice in a fixed MRDF (Modality-Regularization-based DeepFake)-based pipeline with AV-HuBERT visual processing and an MSCT (multi-scale cross-modal transformer) fusion encoder. We compare its inherited spectrogram-style log-filterbank frontend with frozen wav2vec\~2.0 and WavLM representations, dual-SSL fusion, actually executed SSL depth, and shallow-layer aggregation. On the generator-disjoint MMDF dataset, full wav2vec\~2.0 improves mean macro AUROC from $84.19\%$ to $87.19\%$. True six-layer execution improves over the full encoder for all three matched seeds, and learned non-uniform aggregation adds a further $2.53$ points, reaching $90.61\%$ macro AUROC. Matched profiling shows lower computation than full-depth wav2vec\~2.0 across input durations, while a post-hoc first-second evaluation retains $86.85\%$ macro AUROC. A matched dual-SSL diagnostic finds that concatenation outperforms either single SSL stream, suggesting complementary information, whereas one-way cross-attention loses this gain and adds cost. Inference interventions indicate that the aggregation benefit comes mainly from a front-loaded layer mixture rather than strong visual conditioning.
Chat is not available.
Successful Page Load