Which Pairs Should We Compare for DPO?
Abstract
Direct Preference Optimization (DPO) has become a standard method for preference-based fine-tuning in RLHF pipelines. A practical but largely heuristic step in DPO is dataset curation: given a fixed budget of preference labels, which candidate pairs should be compared to produce the most informative training set and the best downstream policy? This paper studies DPO dataset curation as a sampling-design problem. We model the curated dataset by a sampling design over candidate pairs and analyze how this design propagates through DPO training to the KL-regularized RLHF objective. Our main results provide matching upper and lower bounds on the RLHF optimality gap of the policy learned from a curated DPO dataset. The bounds reveal a simple structural message: the effect of pair selection enters only through a single design-dependent matrix that summarizes how informative the collected comparisons are for learning the policy parameters. This leads to a trace-form criterion that characterizes both an achievable performance guarantee for the DPO estimator and an information-theoretic lower bound for any induced policy estimator, thereby identifying a canonical objective for comparison curation under budget constraints.