Prompt-Conditioned Semantic Bottleneck for Cross-Domain Face Attack Detection
Abstract
Learning visual representations that generalize across domains remains challenging when source-domain supervision is entangled with spurious dataset-specific cues. This issue is particularly pronounced in face attack detection (FAD), where attack artifacts vary substantially across capture devices, manipulation pipelines, and image quality. In this paper, we propose Prompt-Conditioned Semantic Bottleneck, an information-bottleneck-inspired representation learning framework for cross-domain FAD. Unlike prior methods that use prompts as image-text matching anchors for multi-modal feature learning, our method uses label- and attack-type-conditioned prompts to elicit attack-aware semantic targets from a frozen VLM. These prompt-conditioned semantics provide offline supervision for a lightweight visual detector, encouraging its representation to preserve attack-relevant cues while reducing reliance on source-domain nuisance factors. Extensive experiments on face anti-spoofing and face forgery detection benchmarks demonstrate consistent improvements under cross-dataset evaluation. Ablation studies further show that the gains are not simply due to stronger visual or multi-modal models, but arise from attack-aware prompt-conditioned semantic supervision. Importantly, the VLM and prompts are removed during inference, enabling efficient vision-only deployment. Overall, our results suggest that prompt-conditioned VLM semantics provide an effective way to improve the trade-off among detection accuracy, cross-domain generalization, and inference efficiency in FAD.