Understanding and Enhancing Conditioning Robustness of Text to Image Diffusion Models
Valery Parfenov ⋅ Timofey Belinsky ⋅ Viacheslav Vasilev ⋅ Denis Dimitrov ⋅ Aleksandr Beznosikov
Abstract
Recent advances in diffusion models have significantly improved the quality of text-to-image generation. Yet visual fidelity alone is not sufficient: reliable conditioning is essential for predictable control, and its limitations continue to constrain the broader success of these models. In particular, the sensitivity of generated images to non-semantic perturbations of the input prompt is a widely recognized problem. In this work, we develop a simple perturbative model of the mechanism underlying this instability and validate it through targeted experiments. We further introduce dedicated metrics and a benchmark (CoRBen) for systematically evaluating prompt robustness. The resulting theory and empirical findings motivate the $\varepsilon^{*}$ sensitivity measure and a family of training-free plug-and-play methods that improve generation stability at inference overheads of 100\%, 17\%, and 3\%. Our code is available at https://anonymous.4open.science/r/conditioning-robustness/.
Chat is not available.
Successful Page Load