SPECTRA: A Privacy-Preserving Graph Architecture for Real-Time On-Device Deepfake Audio Detection
Abstract
The rapid rise of AI-generated voice deepfakes has facilitated mass misinformation and fraud, disproportionately affecting vulnerable populations such as the elderly and individuals with limited digital literacy who rely heavily on mobile communication. Existing detection systems are either computationally intensive and cloud-dependent or insufficiently robust for deployment in resource-constrained environments. To address this gap, this paper introduces Spoofing Protection via Enhanced Contextual Temporal Representation Architecture (SPECTRA), a lightweight deepfake detection framework that combines efficient self-supervised learning feature extraction with a graph-based neural network architecture. It further incorporates GATv2 layers and DropEdge regularization to enhance robustness under noisy conditions. To enable on-device deployment, the system employs a compact backend to reduce computational overhead and a quantized frontend to reduce deployment size. Evaluated on the ASVspoof 2021 DeepFake dataset, a widely used benchmark for real-world spoofing robustness, the full-precision SPECTRA framework achieves an Equal Error Rate of 11.38%, outperforming several strong baselines. Additionally, it operates with a Real-Time Factor of 0.0767 natively on physical mobile hardware and maintains a compact deployment footprint suitable for mobile applications. These results demonstrate that accurate and privacy-preserving deepfake detection can be achieved directly on devices without reliance on cloud infrastructure, thereby supporting scalable deployment in low-resource and mobile-first environments.