Measuring additive and nonlinear concept structure in text-conditioning embeddings with supervised sparse decoders
Abstract
Many interpretability and control methods assume that semantic concepts correspond to reusable directions that compose additively in representation space, motivated by the linear representation hypothesis. We directly test this assumption on the prompt embeddings of Stable Diffusion 3.5 (SD3.5) using a controlled, combinatorial concept dictionary. We formulate additive concept reconstruction as a supervised sparse decoder, providing an interpretable representation in which each concept occupies a predefined latent block. We then extend this formulation with nonlinear interactions, enabling a controlled comparison between additive and non-additive representation structure. We find that SD3.5 text-conditioning embeddings exhibit substantial but incomplete additivity: a simple additive model explains up to 79\% of held-out variance, while nonlinear interactions capture additional predictable structure. The learned representations also generalize to unseen concept combinations and support modular feature-level replacements, while exposing a failure mode of concept deletion. Our methodology thus provides a quantitative framework for testing when concept representations compose and when nonlinear interactions between concepts are necessary.