MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation
Abstract
End-to-end full-duplex speech models have brought open-source machine conversation close to human fluency, yet existing systems fall short of real-world deployment in two entangled respects: long-context conversational robustness and multi-party interaction capability. Realistic settings, including meetings, group lessons, family dinners, and social-robot reception, are inherently long-horizon and multi-party at the same time, requiring a single model to perceive, attribute, contextualize, and respond across multiple speakers over extended durations. Progress along these axes is bottlenecked by both data and evaluation. On the data side, open multi-party conversational speech corpora total only a few hundred hours and are not designed for codec-frame-level full-duplex modeling. On the evaluation side, existing long-audio benchmarks focus on passive listening, while existing speech-to-speech benchmarks remain dyadic and short. In this work, we extend the Moshi paradigm along long-horizon and multi-party axes simultaneously, in both English and Chinese, with three contributions. First, we release an open data engine and a 57.6 k-hour corpus for long, multi-party, bilingual full-duplex dialogue. The engine produces parallel-stream audio with controllable length, participant count, conversational dynamics (turn-taking, overlap, backchannels, interruption, addressee shifts, long-range co-reference), and language (English, Chinese, and intra-sentential code-switching), exceeding all prior open multi-party conversational speech corpora by more than an order of magnitude. Second, we introduce MultiTalkBench, the first benchmark to jointly evaluate long-form full-duplex dialogue with an average duration of 32.6 minutes, multi-party, and bilingual full-duplex dialogue, with explicit probes for long-range entity tracking, topic coherence, and addressee selection. Third, we train a bilingual Moshi-style full-duplex model on the released corpus that sustains coherent multi-party English-Chinese conversation over extended durations, substantially outperforming open-source baselines including Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct.