StableAvatar: Ultra-Long Audio-Driven Avatar Video Generation
Abstract
Current diffusion models for audio-driven avatar video generation struggle to synthesize long videos with natural audio synchronization and identity consistency. This paper presents StableAvatar, the first end-to-end video diffusion transformer that synthesizes ultra-long videos. We observe that the primary reason preventing existing models from generating long videos is their audio modeling, which typically relies on third-party off-the-shelf audio embeddings that are introduced into the diffusion model via cross-attention. Since current diffusion backbones lack any audio-related priors, they cannot correctly couple audio cues with latent dynamics, causing systematic latent distribution deviations at each segment. These deviations accumulate across segments, gradually shifting the latent trajectory and producing noticeable quality drift. To address this, StableAvatar introduces a novel Timestep-aware Audio Adapter that prevents error accumulation via timestep-aware modulation. During inference, we propose a novel Audio Native Guidance Mechanism to further enhance the audio synchronization by leveraging the diffusion’s joint audio-latent prediction as a dynamic guidance signal. To enhance the smoothness of the ultra-long videos, we introduce a Dynamic Weighted Sliding-window Strategy that fuses latent over time. Experiments on benchmarks show the effectiveness of StableAvatar both qualitatively and quantitatively.