Architecture Shapes Emergent Full-Duplex Speech in Audio-Video Diffusion Models
JIA-HONG HUANG ⋅ Jiajun Fan ⋅ Shuai Wang ⋅ Ge Liu ⋅ Prayag Tiwari ⋅ Stevan Rudinac
Abstract
Human conversation is inherently full-duplex, with speakers naturally producing interruptions, backchannels, and overlapping speech. However, most speech generation systems remain fundamentally turn-based and require explicit mechanisms to model duplex interaction. In this work, we investigate whether full-duplex conversational behavior can emerge naturally in general-purpose multimodal generative models without explicit overlap supervision. We discover that LTX-2, a large-scale diffusion transformer trained for synchronized video and audio generation, exhibits emergent full-duplex conversational speech generation despite having no duplex-specific training objective. Through systematic evaluation across five levels of overlap complexity, we show that LTX-2 generates coherent multi-speaker conversations with increasing overlap as conversational complexity grows, while remaining constrained by the structure of natural audiovisual interactions. To understand this emergence, we propose the visual grounding hypothesis, which argues that spatial, temporal, and scene-level visual information provides structural organization for multi-speaker interaction, enabling overlapping speech when audiovisual and conversational cues are aligned. Further analysis reveals that the emergent behavior is controllable but fragile: LoRA adaptation amplifies targeted overlap patterns by up to 5.5$\times$, while narrow fine-tuning can compromise broader multi-speaker capabilities. Cross-architecture analysis with MOVA further shows that audiovisual co-generation alone is insufficient to induce duplex behavior, highlighting persistent cross-modal coupling within a shared backbone as an important architectural factor. Our findings reveal that multimodal generative models can acquire interaction capabilities beyond their explicit training objectives and provide insights into how visual grounding, architecture, and adaptation jointly shape emergent conversational behaviors. We will release the codebase publicly upon acceptance to facilitate reproducibility.
Chat is not available.
Successful Page Load