Tracing the Geometry of Compositional Generalization in a Diffusion U-Net
Abstract
Diffusion models can perform compositional generalization, creating unseen combinations of familiar features, but how this ability is acquired internally remains unclear. We study the transition from memorization to compositional generalization in a conditional U-Net trained on a diffusion task with images of three binary factors. Representational changes are concentrated in the decoder, where feature geometry reorganizes and representation dimensionality expands, while encoder representations remain comparatively stable. Causally, transplanting post-generalization decoder blocks Up1+Up2 weights into the pre-generalization model recovers 93\% of the out-of-distribution performance gap. Steering learned feature directions also selectively controls generated attributes. These results identify localized decoder changes associated with the emergence of compositional generalization, advancing our understanding of how generative models acquire generalizing behavior.