Past as State, Present as Attention: Persistent-State Blockwise Flow Matching for Long-Horizon Co-Speech Motion Generation
Yi Liu ⋅ Xiangyue Zhang ⋅ Jia Ma ⋅ Jianfang Li ⋅ Yisheng He ⋅ Jianqiang Ren
Abstract
Co-speech motion generation remains challenging over long horizons due to the need for temporal coherence and consistent style. Existing methods typically generate long motions segment by segment, either conditioning on a short seed of recent frames, which gradually drifts, or retaining ever-longer explicit histories, which are costly to maintain and hard to learn from limited data. We trace these limitations to a key asymmetry between past and present: past motion carries slowly varying style that should be compactly summarized, whereas the current segment exhibits fast-changing local dynamics that require fine-grained modeling. This asymmetry motivates a simple $\textbf{P}$ast-$\textbf{A}$s-$\textbf{S}$tate, $\textbf{P}$resent-as-$\textbf{A}$ttention principle for long-horizon co-speech motion generation. Following this principle, we propose $\textbf{PASPA}$, which summarizes past history with an inter-block persistent state and models the current segment with intra-block bidirectional self-attention. We instantiate PASPA within a blockwise flow matching framework, with multi-block supervision for effective training over variable histories, prefill-decode state caching for efficient rollout, and hybrid classifier-free guidance for enhanced history and speech conditioning. Extensive public-benchmark experiments show that PASPA achieves state-of-the-art FGD performance and preserves motion style across long-sequence generation.
Chat is not available.
Successful Page Load