MUNI: Multimodal Unified Latent Diffusion for Coherent Any-to-Any Generation
Kyeongmin Yeo ⋅ Yunhong Min ⋅ Minhyuk Sung
Abstract
We introduce $\texttt{MUNI}$, an end-to-end multimodal latent diffusion framework for any-to-any generation that unifies subset-conditioned cross-modal generation and unconditional joint sampling through a shared stochastic latent. Existing multimodal generative models are largely LLM-based, which limits leveraging modality-specific generators and requires text-paired data for training. Recent diffusion- and flow-based any-to-any extensions take a different direction but still rely on text-aligned embeddings, fully-paired training, or matched-dimensionality deterministic mappings. $\texttt{MUNI}$ rests on two complementary contributions, one architectural and one in the training objective. First, we extend latent diffusion to multimodal any-to-any generation *end-to-end*: instead of the standard two-stage recipe that precomputes a frozen latent space and then fits a prior over it, $\texttt{MUNI}$ jointly trains modality-specific encoders, expressive decoders, and a single shared flow-based prior under one objective. Second, we identify that the standard aggregation rules of multimodal variational inference are insufficient once coupled with a learned prior and expressive decoders. A suitable shared latent must simultaneously satisfy *coherence* across generated modalities, *predictive sufficiency* of subset latents, and *minimality* of the latent content. We propose a routed training objective whose structural choices align the latent with these criteria and admit a minimal-sufficiency characterization in the realizable setting. Experiments on PolyMNIST-Quadrant-Labels and a large-scale image-text-audio benchmark show $\texttt{MUNI}$ matching or exceeding the strongest baselines on conditional generation while opening its largest margins on unconditional coherence.
Chat is not available.
Successful Page Load