Predicting Protection Strength for Diffusion-Based Image Editing Without the Attacker's Model
Saarthak Gupta ⋅ Vasu Mahajan ⋅ Arjun K Tomar ⋅ Vedant K Shanker
Abstract
Protective perturbations defend images against unauthorized diffusion-based editing, but a protector must commit to a perturbation before knowing which editing model an adversary will use. We introduce the specificity ratio $\sigma$, which measures how efficiently a perturbation displaces a fixed CLIP image representation relative to random noise of the same magnitude. It is computed from forward passes through that one encoder and requires no access to any unseen model. Across nine perturbation instruments, 20 images and five unseen diffusion models, we find $\sigma$ predicts how much a perturbation disrupts editing: within-image Spearman $\rho = +0.87$ on held-out instructions and $+0.64$ on unseen models, with every image and 95\% of images agreeing in sign respectively. The relationship holds after controlling for perceptibility and replicates across two independent runs. We also report a methodological finding that we think is the more useful contribution. Following common practice, we initially normalized transfer by in-domain strength and observed an apparent strength--transferability trade-off ($\rho = -0.45$). Because $\sigma$ also correlates with in-domain strength, this ratio can produce a negative correlation mechanically. Three analyses that avoid the division, a partial correlation, a per-image regression, and a residual correlation, all reverse the sign to $+0.41$, $t = +4.26$, and $+0.48$ respectively. The apparent trade-off was an artifact of the denominator. Since ratio-normalized transfer metrics are widespread in this literature, we document the failure mode in full.
Chat is not available.
Successful Page Load