RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space
Abstract
Large language models (LLMs) are widely used in text-to-image (T2I) systems, but they are typically limited to text encoding, while denoising is handled by newly trained generative backbones. The emergence of representation autoencoders (RAEs) shifts the generation target toward semantically structured visual representations, creating a latent space that is more compatible with pretrained LLM priors. Inspired by multimodal LLMs (MLLMs), where an MLP projector is sufficient to align clean visual representations with a pretrained LLM, we repurpose the MLLM itself as a noisy representation encoder, extending this mechanism from clean to noisy inputs. We present RepFusion, which uses the resulting MLLM outputs as the conditioning signal for a diffusion transformer. In controlled, parameter-matched comparisons, RepFusion outperforms scaling the denoising transformer with newly initialized parameters. These results demonstrate that MLLMs provide strong priors for denoising visual representations and that scaling the compute of the conditional encoder is feasible in modern T2I systems.