Real-Time Conversational Agents: Toward Natural Multimodal Interaction
Abstract
Real-time conversational agents have rapidly moved from research demonstration to deployed product, with voice modes, embodied avatars, and full-duplex speech systems now powering applications across customer service, education, healthcare, and accessibility. To feel natural, however, such systems cannot rely on the offline generation paradigm that has dominated multimodal machine learning: they must produce speech, video, and language in a streaming fashion while continuously listening, watching, and re-planning over partial observations. The methodological constraints introduced by this regime, namely sub-second latency budgets, causal and incremental computation, full-duplex audio modelling, tight cross-modal temporal alignment, and handling of interruptions, backchannels, and overlapping speech, differ qualitatively from those addressed by the offline literature, and techniques that have driven offline progress (non-causal attention, large-beam decoding, multi-pass refinement, slow diffusion sampling) frequently fail to transfer. Despite recent advances in full-duplex audio–language modelling, real-time talking-head and avatar synthesis, low-latency speech generation, and streaming automatic speech recognition, deployed agents remain perceptually robotic: turn-taking is stilted, backchannels are absent, prosody is monotone, and gaze and gesture are routinely mis-timed with respect to the linguistic and affective content of the utterance. The community has not yet converged on a shared vocabulary, benchmarks, or methodology for evaluating interactional naturalness as distinct from per-utterance quality measured by mean opinion scores or offline win-rates. The Workshop on Real-Time Conversational Agents (RTCA) addresses this gap by convening researchers across speech, vision, language, human–computer interaction, social signal processing, and machine learning systems around three intertwined questions. First, real-time generation: how to produce high-quality speech, video, and language under hard latency budgets, in a streaming or full-duplex fashion, with attention to the architectural, training, and inference-system innovations that distinguish streaming from offline modelling. Second, naturalness in interaction: what perceptual, linguistic, and behavioural elements prosody, gaze, timing, grounding and expressivity. Third, evaluation of live systems: how to design metrics, benchmarks, and study protocols that capture naturalness, responsiveness, and conversational quality in interactive settings, where standard offline metrics and held-out test sets are demonstrably inadequate. The workshop solicits short papers, full-papers, and demo papers on streaming speech synthesis and recognition, full-duplex audio–language models, real-time talking-head and embodied avatar generation, incremental and speculative decoding, turn-taking and floor management, multimodal alignment under partial observation, prosody and paralinguistic generation, memory and grounding in live conversation, interactive evaluation protocols, efficient inference and on-device deployment, and the safety and identity considerations specific to real-time multimodal generative agents. The programme combines invited talks across the workshop's four thematic pillars with contributed talks and posters, a live Conversational Agents Showcase in which accepted demonstration systems are run on stage, and a closing panel on what naturalness in conversational AI actually means and how it should be measured. By foregrounding evaluation alongside generation and by convening communities that currently publish across separate venues, the workshop aims to consolidate shared problems, datasets, and methodological standards for a research area whose downstream user impact already substantially exceeds its academic visibility.