UNICOM: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
Abstract
Current unified multimodal models typically rely on discrete visual tokenizers to bridge the modality gap. However, discretization inevitably discards fine-grained semantic information, leading to suboptimal performance in visual understanding tasks. Conversely, directly modeling continuous semantic representations (e.g., CLIP, SigLIP) poses significant challenges in high-dimensional generative modeling, resulting in slow convergence and training instability. To resolve this dilemma, we introduce UniCom, a unified framework that harmonizes multimodal understanding and generation via compressed continuous semantic representations. We empirically demonstrate that channel-wise compression is significantly more effective than spatial token reduction, retaining competitive VAE-free reconstruction fidelity while substantially improving generative convergence. Accordingly, we design an attention-based semantic compressor to distill dense features into a compact unified representation. A Transfusion-style predictor then generates these compressed latents with dense spatial correspondence. At scale, UniCom achieves competitive text-to-image generation and state-of-the-art performance on complex image editing without relying on VAE latents. The gains are especially clear on knowledge-intensive benchmarks.