SitCom: Scaling Egocentric Multi-Party Spoken Dialogue for Situated Communication Assistance
Abstract
Multi-party spoken communication pervades daily life, yet even attentive participants routinely lose track of who said what, miss a turn, or struggle to recall something said minutes earlier. As always-on wearables become commonplace, an assistant that listens alongside the user could ease this everyday friction. We introduce \textit{Situated Communication Assistance}: on-demand support for users immersed in ongoing multi-party conversations, grounded in the same egocentric acoustic scenes they perceive. We argue this capability rests on three abilities: producing a Rich Situated Transcription (RST) of \textit{who} said \textit{what}, \textit{when}, and \textit{where}; comprehending the conversation holistically; and answering user queries while the conversation is unfolding. To address the persistent data deficit in this regime, we build \textbf{SitCom}, a multi-party spoken dialogue corpus more than an order of magnitude larger than the largest prior multi-party corpus, pairing 15.8k hours of synthesized data with 144 hours of real recordings standardized under a unified RST schema. The synthesis pipeline generates group-structured scripts with persona-grounded speech and behaviors, and renders them in 3D scenes through a directional simulator with head-and-torso radiation. Using this corpus, we curate \textbf{SitCom-Bench}, a fully human-validated question answering benchmark of 4.3k items spanning seven comprehension and four in-the-moment query subtypes. Experiments show that combining synthetic and real data is necessary to close the sim-to-real gap, and that in-the-moment querying remains hard even for frontier models, pointing to it as an open challenge distinct from post-hoc comprehension.