Mult-DPO: Multinomial Direct Preference Optimization for Recommender Systems
Abstract
Direct preference optimization (DPO) is a simple and effective alignment strategy for large language models (LLMs) based on pairwise preferences between two candidates. In recommender systems, however, user feedback is rarely pairwise. For a given context, e.g., a user, a session, or a conversation, we typically observe set-wise preferences with multiple positive items, where every positive item should outrank every unobserved or explicitly negative item, with no prescribed order among the positives or the negatives themselves. A natural generalization is to use the Plackett–Luce (PL) reward model, which extends the Bradley–Terry (BT) reward model underlying vanilla DPO from pairwise preferences to full rankings of candidates. However, adapting the PL model to set-wise preferences requires marginalizing over all consistent permutations of the positives and negatives respectively, which is intractable. To address this fundamental challenge, we propose Mult-DPO, a novel DPO objective with a tractable multinomial reward model over set-wise preferences. We show that, like the BT and PL reward models, the multinomial reward admits a closed-form DPO-style objective, enabling direct alignment of LLMs without reinforcement learning (RL). In addition, we prove that the multinomial DPO loss is a tractable upper bound on the exact marginalized PL DPO loss when optimizing against the set-wise preference data. We further characterize the tightness of this bound in terms of relative total weight (i.e., exponentiated rewards) of positives versus negatives, which provides insights into tightening the bounds with more or harder negatives. Finally, we extend the framework to a group-wise setting that accommodates multiple preference levels. Code and datasets are available at https://anonymous.4open.science/r/mult-dpo-41C7.