Tap2Listen: Instant Target Switching for Streaming Audio-Visual Speaker Extraction
Hansol Park ⋅ Minkyung Song ⋅ Zinu Kim ⋅ Yu J Lee ⋅ Hoseong Ahn ⋅ Hyunjin Choi ⋅ Kyuhong Shim
Abstract
Streaming audio-visual target speaker extraction (AV-TSE) uses a target speaker's face to extract the speaker's speech from an input audio mixture in real time. However, existing systems fix the target speaker for an entire clip, while conversation demands frequent and immediate target switches. To address this problem, we propose Tap2Listen, a streaming AV-TSE model designed for instant response to target switches. The model separates computation by modality: the target-agnostic audio encoder processes only the mixture, while the target-specific visual cues of the selected face condition only the decoder. A switch requires only a $0.12$ ms re-projection and takes effect immediately. We also design component-wise losses: a prediction loss for the encoder and a supervised extraction loss for the decoder. For evaluation, we measure how quickly output quality recovers after a switch. On the standard LRS3-2mix test set, Tap2Listen improves SI-SNRi from $8.85$ to $9.68$ dB at a comparable parameter count. After a switch to a known face, the model produces full-quality output from the first frame, whereas the baseline needs about 1 second to recover. Finally, we show that the proposed streaming AV-TSE model runs in real time on a single CPU core.
Chat is not available.
Successful Page Load