TurnMoshi: Pruning and Distillation of a Full-Duplex Model for Real-Time Turn-Taking
Abstract
Turn-taking is a core capability of spoken dialogue systems, requiring models to continuously coordinate conversational dynamics. Recent full-duplex dialogue models learn these behaviours directly from speech, raising the question of whether their internal representations can be reused for turn-taking without retaining the full capacity required for speech understanding and generation. We investigate this using Moshi, a full-duplex speech model. We probe its intermediate representations with Voice Activity Projection (VAP) to identify separate layers that encode useful information for end-of-turn and interruption prediction. This is followed by layer truncation and structured pruning with supervised and teacher-student distillation, while keeping the pretrained Moshi backbone frozen. On the TurnBench development set, TurnMoshi reaches 0.908 recall for end-of-turn prediction compared with 0.836 for a VAP baseline, and matches VAP on interruption prediction. We find that this approach is also robust to noise (retains 90% of its average clean EOT/INT recall under unseen noise at -10 dB SNR, compared with 72% for VAP). TurnMoshi also continues to outperform VAP on the Switchboard corpus. These results indicate that turn-taking representations learned within a generative full-duplex dialogue model can be isolated and repurposed for dedicated real-time turn-taking prediction.