StreamSV: Streaming Speaker Verification for User-Specific Interaction with Voice Assistants
Abstract
Conventional speaker verification systems operate on pre-segmented utterances of fixed duration, limiting their applicability in real-time and interactive scenarios. These limitations are especially relevant to full-duplex voice assistants, where recognizing the intended user is essential for natural, user-specific interaction. One example is personal interruption, where only the enrolled user should be able to interrupt the assistant’s response. To address this problem, we present StreamSV, a streaming speaker verification system for short utterances that combines a fully causal embedding extractor with a lightweight recurrent scoring head. The encoder is trained directly on speech prefixes, and the head turns their embeddings into a continuously updated probability of speaker match. Verification scores are produced every 100\,ms, with each incoming chunk processed in less than its own duration. This supports real-time operation and adaptive stopping, allowing clear trials to terminate early while ambiguous trials accumulate more speech. Experiments on VoxCeleb1-O demonstrate strong short-utterance performance, reaching 6.41\% equal error rate at 0.5\,s, while adaptive stopping achieves 1.85\% overall error with an average decision time of 847\,ms. Ablation studies confirm that both encoder adaptation and recurrent scoring contribute to these gains. These properties support the use of StreamSV as a component for user-specific interaction with full-duplex voice assistants.