Predicting Human-Gain Curves under Local Proxy Optimization from Repeated Ratings
Moonwon Choi ⋅ Kisung Nam ⋅ Seunggeun Lee
Abstract
Optimizing a proxy objective can improve human-perceived quality, but the resulting gain can vary with optimization strength. Average proxy--human correlation alone does not reveal how human-perceived quality changes as optimization strength increases. We propose a pre-optimization diagnostic that uses repeated human ratings to predict the human-gain curve induced by a fixed proxy. The method models proxy optimization as a KL-regularized exponential tilt. At KL radius $d$, the local gain is approximated by $A\sqrt{d}+Bd$, where $A$ measures first-order proxy--human alignment and $B$ captures the leading second-order correction along the proxy-optimization path. We estimate $A$ and $B$ from repeated human ratings on calibration items and test the resulting gain-curve prediction on held-out evaluation items. Across SummEval, analysis-eligible WMT MQM clusters, an OpenMEVA-MAGS ROC/WP repeated-rating slice, and a GPT-4o judge proxy experiment, the proposed approach consistently reduces held-out gain-prediction error relative to $A$-only, zero-gain, and reliability-only baselines. Transfer experiments across datasets and evaluation slices show that source coefficients are useful priors, but target calibration further improves prediction. The method provides a principled local diagnostic for estimating the human-gain curve of a fixed proxy as optimization strength varies.
Chat is not available.
Successful Page Load