Scalable Minimal-Change Learning for Controllable Image Editing
Shuo Chen ⋅ Fengming Huang ⋅ Yu Yao ⋅ Mingming Gong ⋅ Tongliang Liu
Abstract
Image editing aims to modify specific attributes of an image while preserving all other aspects. In practice, however, applying even simple editing instructions to existing methods often leads to unintended changes. We identify the central issue: a fundamental causal principle of \emph{minimal change}, which requires that an intervention alter only the intended attributes in the output while leaving all others invariant, has not been explicitly incorporated as an optimization objective. To encourage minimal change, prior work rooted in causal representation learning typically imposes an $L_1$ regularizer on the latent difference between pre- and post-edit representations. However, this strategy does not scale to modern image editing models: $L_1$ regularization fails to induce true sparsity and instead tends to shrink all latent differences uniformly; sparsity in the latent space does not translate to localized change in the output of nonlinear deep image generation models; and the multi-domain or counterfactual supervision it requires is generally unavailable. To effectively leverage the minimal change principle for controllable image editing, we cast it as a reinforcement learning objective. Rather than constraining latent representations, we design rewards that directly capture minimal change over intervention outcomes and optimize them explicitly during training. To produce these rewards reliably, we further introduce an agentic reward model that leverages the chain-of-thought reasoning ability of vision-language models without requiring human annotations. The idea is that instead of emitting a single, unreliable scalar score, the model reasons explicitly about two complementary failure modes, unimplemented changes and unintended changes, and produces this supervision automatically. Experimental results demonstrate that our method produces significantly more localized and semantically consistent edits compared to existing approaches, reducing unintended changes.
Chat is not available.
Successful Page Load