Testing and Estimation of Contextual Generalized Thurstone Models
Dongmin Lee ⋅ Anuran Makur ⋅ Japneet Singh ⋅ Boyu Xu
Abstract
Many modern machine learning methods, such as reinforcement learning from human feedback (RLHF) to train reward functions, utilize preference learning models like Bradley-Terry-Luce (BTL) to learn latent scores of items based on pairwise comparison data. Despite the successes of such models in specific applications, the theoretical justification for the employed models has often been unclear. In this work, we provide a hypothesis testing procedure to test whether a given dataset of pairwise comparisons follows a contextual generalized Thurstone (CGT) model under certain classes of latent functions. The latent functions determine the latent scores, and can be expressed as weighted sums of basis functions. Our CGT model encompasses a wide class of preference learning models, including the popular BTL model. In particular, we prove critical thresholds for our tests in the minimax sense, showing that they scale like $1/\smash{\sqrt{|\mathcal{E}|^{1/2}k}}$ under certain important regimes, such as complete graphs and perfect matchings. Furthermore, to establish our testing results, we also derive error bounds for CGT parameter estimation that generalize and improve on known results in the BTL and non-contextual settings. Finally, we conduct experiments on various synthetic and real-world datasets to verify our theoretical results.
Chat is not available.
Successful Page Load