Where Evaluators Agree: Quality-Dependent Agreement in Text-to-Image Reward Optimization
Abstract
Reward models (RMs) are widely used to optimize generative models, yet distinct evaluators can rank the same outputs differently. We study how evaluator agreement varies across reference-ranked quality regions in text-to-image generation. Using four public image-preference RMs and a cross-fitted analysis of HPDv2 over all 252 possible reference--evaluation annotator assignments, we find that RM–RM, RM–human, and human–human pairwise agreement are consistently higher among bottom-ranked outputs than among top-ranked outputs. Standardizing the two regions to the same distribution of reference-group vote margins attenuates but does not eliminate these gaps, indicating that pair difficulty explains part, but not all, of the observed asymmetry. Motivated by this structure, we evaluate two single-RM interventions that restrict optimization pressure: within-group quantile clipping and a capped lower-tail objective. In a single-seed Stable Diffusion v1.5 study, Quantile Clip yields higher point estimates than group-normalized DDPO on all three non-optimized RMs, while Capped CVaR does so on two of three; naive CVaR improves none. These results provide preliminary evidence that the location of reward-optimization pressure can affect cross-evaluator transfer.