Real-Time Multimodal Conversational AI
Abstract
For a long time, conversational AI has been confined to disembodied and unimodal text or speech exchanges. As we enter the era of embodied agents and virtual assistants, human-machine dialogue is increasingly being rooted into the physical world. This transition requires agents, whether virtual or physically embodied, to perceive scenes and humans through multimodal signals (e.g. speech, video, sensor recordings, etc.), to generate speech and non-verbal communication while maintaining a coherent conversation over time and to act, if necessary, in the physical world. This process must happen in real-time, with low latency and in a contextually-aware manner to enable a seamless human-machine collaboration. Historically, research on these challenges has been fragmented across communities, addressing isolated aspects rather than the problem as a whole. For instance, egocentric conversational AI focuses on the communication with an agent viewing the world as the user sees it. Examples include augmented-reality glasses, AI assistants such as Gemini Live or MiniCPM-o 4.5, or wearable cameras. These systems do not interact directly with the physical world. On the other hand, exocentric and dyadic conversational AI focuses on third-person perspectives and face-to-face communication, including traditional robotics or the more recent Interaction Model. However, real-time human-machine multimodal conversations inherently demand multiple perspectives. A seamless dialogue about a shared task, such as an AI assistant helping a human to assemble furniture, requires understanding both what the human sees (egocentric) and the context of the room, objects (exocentric) and human emotions, gesture and motion (dyadic) to perform cross-view reasoning. These challenges collectively define four foundational research axes requiring interdisciplinary collaboration across computer vision, robotics, machine learning, speech and audio processing and understanding, along with dialogue communities: multimodal representation and perception, real-time interaction dynamics and memory, emboddied interaction and finally, benchmarks and datasets. The 1st Real-Time Multimodal Conversational AI workshop will enable and catalyze progress along these four outstanding research problems.