GESTO: Online Co-Speech Gesture Generation with a One-Step Causal Flow Model
Abstract
Continuous diffusion and flow-matching models produce high-quality co-speech gestures, but they rely on future speech context and multiple sampling steps, which rules them out of low-latency streaming interaction. Existing systems that integrate spoken response and gesture instead represent motion as discrete tokens and predict them autoregressively, leaving a gap between high-quality continuous motion and real-time speech–gesture generation. We present GESTO, a streaming speech-and-gesture framework that couples a spoken-dialogue model with a one-step flow-matching gesture generator: as the response is spoken, GESTO turns the dialogue model's intermediate speech representations into temporally aligned continuous motion latents, one short chunk at a time and without access to later chunks. To make causal generation efficient, we distill a bidirectional flow-matching teacher into a one-step chunk-causal student trained on its own streaming self-rollouts. On the BEAT2 multi-speaker evaluation GESTO reaches a Fréchet Gesture Distance of 0.236, the best among the compared streaming methods, and participants prefer its motion to an autoregressive speech-and-gesture baseline for both quality and speech appropriateness, on prerecorded speech and on responses generated by the dialogue model. The same pipeline trains on retargeted humanoid motion, extending it beyond virtual human animation.