People Making Conversation: Evaluating How Naturally Speech Models Converse
Abstract
A conversational speech model has to decide, at every pause, whether to speak. Existing evaluations do not test this decision. They test whether a system detects a boundary, follows a stated turn-taking policy, or matches human conversation in aggregate. We built a benchmark that tests the decision itself. We take 176 casual conversations from Seamless Interaction, label every turn, pause, and backchannel, and select 427 pauses inside a speaker's turn and 482 turn ends where the right choice is not obvious. A system is given the conversation so far and hears the other speaker in real time. At a pause we measure whether it speaks, and how soon; at a turn end, how fast it answers, how long it talks relative to the person, and how close its words are to theirs; and after 30 to 120 seconds of history, whether it can answer a question about the conversation. We evaluate six real-time systems: PersonaPlex, Moshi, GPT Realtime at two endpointer settings, Gemini 3.1 Flash Live, and Grok Voice. Once its endpointer hears silence, every hosted system speaks after every pause; whether its words land inside the pause is set by its latency, 1 to 4 s, not by a decision. The full-duplex models are the only systems that ever stay silent, and the only ones that interrupt within a human reaction time, 220 to 370 ms. At turn ends all systems talk longer than the measured human response. Given the history as text, the hosted systems answer 33 to 61% of questions about it; hearing it as audio, the full-duplex models answer 2 to 22%. We release the corpus selection, the labels, the evaluation points, the harness, and every system output.