Single-Pass Evidence Measurement for Interpretable and Uncertainty-Aware Multimodal Face Anti-Spoofing
Abstract
Face anti-spoofing (FAS) is critical for protecting face recognition systems against presentation attacks in open environments. Beyond high cross-domain generalization, practical FAS deployments must provide interpretable and uncertainty-aware decisions when multimodal evidence is incomplete, degraded, or conflicting. However, current multimodal domain-generalized FAS methods focus primarily on feature fusion or domain alignment, compressing RGB, depth, and infrared cues into opaque logits. This makes it difficult to discern which modality drives the decision, whether modalities reinforce or suppress each other, and whether a given sample lies within the classifier's reliable operating regime. We introduce SEM-FAS, a quantum-inspired yet fully classical single-pass evidence measurement framework that jointly enhances generalization, interpretability, and reliability. SEM-FAS encodes each modality as a normalized complex-valued state whose amplitude captures evidence strength and whose phase encodes cross-modal interference; multiple coherent heads are then mixed into a density-matrix representation. A constrained POVM-like measurement head converts this representation into calibrated class probabilities while analytically decomposing every prediction into per-modality main effects and pairwise synergy/suppression terms. Simultaneously, three built-in uncertainty indicators are obtained without sampling overhead: predictive ambiguity, representation mixedness, and measurement-coverage mismatch. On standard multimodal DG-FAS benchmarks, SEM-FAS reduces the best baseline HTER by 2.18\% and improves AUC by 1.38\% under both complete and missing-modality protocols. Visual analyses further suggest that the learned modality effects, pairwise interactions, and uncertainty scores align with expected FAS behavior, even for hard samples, indicating that single-pass evidence measurement is a promising way to jointly enhance generalization, interpretability, and reliability in multimodal FAS.