What kills $v$-prediction? A Patch-wise PCA Perspective on Pixel-Space Flow Matching
Tianze Luo ⋅ Haofeng Huang ⋅ Yuan Yao
Abstract
Latent Diffusion Models dominate generative modeling but face two-stage training complexity and the inherent reconstruction limits of pre-trained autoencoders. Pixel-space Flow Matching offers a simplified end-to-end alternative yet often encounters training instabilities when using the standard $v$-prediction objective. Recent research attributes this difficulty to the off-manifold nature of noised targets, prompting a widespread shift toward $x$-prediction. In this paper, we carefully investigate this prevailing assumption through patch-wise Principal Component Analysis (PCA), and we reveal that the failure of $v$-prediction is primarily a bandwidth issue. Specifically, the velocity field in minor components is mostly determined by the raw noisy input, which differs from the dynamics in the principal component space. Deep backbones must propagate this minor information to the final layer to predict the velocity field, which consumes significant model capacity that should be dedicated to principal components. This is particularly critical in pixel space where noise dimensionality is comparable to the hidden dimension. To resolve this, we simply incorporate a term containing the raw noisy input into the final layer. This architectural bypass enables the backbone to focus on high-level semantics while the final layer efficiently reconstructs the minor velocity field. Our method successfully revives $v$-prediction and achieves competitive performance on ImageNet $256\times256$ dataset compared to $x$-prediction baselines.
Chat is not available.
Successful Page Load