SAM-VAP: Semantic-Aware Multimodal Voice Activity Projection for Live Human-Agent Turn-Taking Prediction
Abstract
Turn-taking and listener behaviors such as backchannels are fundamental components of face-to-face interaction, yet remain difficult to predict in a unified multimodal framework. Recent models such as Voice Activity Projection (VAP) achieve strong turn-taking performance from audio and nonverbal cues, but do not leverage semantic information and offer limited support for listener behaviors. We propose a semantic-aware multimodal extension of VAP that integrates speech semantics from a pretrained foundation model (Whisper) through a hierarchical cross-attention architecture, enabling joint prediction of speaker changes and of the listener's verbal and head nod backchannels. We fine-tune the model on the French NoXi corpus, which we extend with automatic annotations of verbal and head nod backchannels derived from linguistic and motion-based criteria. Ablations show that semantic embeddings improve turn-taking prediction, that nonverbal cues are critical for head nod prediction and classification, and that fusing both yields the most balanced performance across tasks. We further deploy the model in an embodied socially interactive agent (SIA), where streaming inference over a reconstructed dyadic audio channel and the participant's live facial features drives the agent's listening behavior in real time, and evaluate it in an experimental study with 60 participants. Participants perceived the fully multimodal configuration as significantly more behaviorally aligned than a rule-based listener, despite its producing markedly fewer feedback actions.