UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
Abstract
Audio generation, including speech, music and sound effects, has advanced rapidly in recent years. These tasks can be divided into two categories: time-aligned (TA) tasks, where each input unit corresponds to a specific segment of the output audio (e.g., phonemes aligned with frames in speech synthesis); and non-time-aligned (NTA) tasks, where such alignment is not available. Since modeling paradigms for the two types are typically different, research on different audio generation tasks has traditionally followed separate trajectories. In this work, we propose UniFlow-Audio, a non-autoregressive audio generation framework that unifies the two types based on flow matching. UniFlow-Audio introduces a dual-fusion mechanism that aligns audio latents with TA features while integrating NTA conditions via cross-attention in each block, and employs task-balanced sampling to maintain consistent performance across diverse tasks. Supporting omni-modal inputs including text, audio, and video, UniFlow-Audio achieves strong results on 7 audio generation tasks with fewer than 8K hours of public training data and under 1B trainable parameters. A compact 200M-parameter variant remains competitive, indicating UniFlow-Audio as a promising foundation model for general non-autoregressive audio generation. Codes and models are available at https://anonymous3387a8c.github.io/uniflow_audio.