Stability-Weighted Direction Regularization Disentangles Generator Shortcuts from Detection Signal
Abstract
AI-generated image detectors often generalize poorly to unseen generators because the classifier latches onto generator-specific feature directions that are predictive only on the training set. We show that simply projecting away generator-discriminative LDA directions fails: these directions also carry genuine real-vs-fake signal, so post-hoc erasure removes evidence along with shortcuts. We introduce StaR, a training-time regularizer that scores each LDA direction by leave-one-generator-out stability and penalizes the classifier only on unstable directions; stable directions are left available for classification. On ResNet-50, StaR reduces the across-seed standard deviation of the four-OOD average AUC by 24x relative to ERM (F-significant on every OOD) while preserving in-distribution-like performance; on CLIP ViT-L/14, StaR improves the four-OOD average AUC by 1.7 points over matched CLIP-ERM and 4.5 over EFFORT, with the largest gains on platform-shift benchmarks. Mechanistic analyses show that StaR rotates the classifier away from unstable generator directions while amplifying generator-discriminative structure in the features---a rotational, not erasive, effect that post-hoc projection cannot reproduce. We additionally find that the conventional argmax in-domain validation'' epoch-selection rule is sub-optimal: a deployablestop at first in-domain 0.99'' rule beats it by sim0.7 AUC points on average.