One Unified Representation: Resolving the Appearance-Semantics Dilemma via Structural Regularization
Abstract
The Representation Autoencoder (RAE) shows that a frozen high-dimensional semantic ViT representation enables both image reconstruction and DiT-based generation. However, RAE's semantic-only representation lacks appearance details, yielding suboptimal reconstruction. Full fine-tuning the encoder with reconstruction loss induces catastrophic semantic collapse: the encoder overfits to appearance, loses semantic discriminability, and cripples generation. This failure arises because the reconstruction loss imposes an appearance prior that erodes semantic structure. To resolve this conflict, we propose Semantic Structure Regularization (SSR). SSR builds on self-distillation, using a learnable student and a frozen teacher that retains the original semantics. In addition to standard latent alignment, which forces the student's representation to match the teacher's, we introduce two complementary regularizers that prevent appearance overfitting from corrupting semantic structure. First, spatial self-similarity distillation forces the student to reproduce the teacher's patch-wise relational patterns. Invariant to appearance changes, this relational signature anchors high-order semantics and prevents the sacrifice of structure for pixel fidelity. Second, structure-anchored alignment decouples structure from appearance. Using a frequency-domain transform, we extract a high-frequency structural map (edges, boundaries) from each image while discarding low-frequency appearance. The teacher's representation of this structural map then regularizes the student's representation of the original image. Consequently, the student learns appearance details from the reconstruction loss without penalty, where the regularization enforces only structural consistency and enables pixel-level encoding without semantic drift. The result is a unified representation that excels at high-fidelity reconstruction, discriminative perception, and efficient DiT-based generation. Our work provides a principled path toward a single semantic space bridging perception and generation, challenging the classic dichotomy. Code and models will be released.