Robust Nash Alignment under Preference Uncertainty
Shihab Ahmed ⋅ Debamita Ghosh ⋅ David Tang ⋅ Yudan Wang ⋅ Alvaro Velasquez ⋅ Yue Wang
Abstract
Preference-based alignment methods typically optimize against a single preference model, and can therefore be brittle when pairwise preferences are uncertain: noisy, heterogeneous, or shift after deployment. We study alignment under uncertain preferences through the lens of general-preference games. Specifically, we formulate Robust Nash Learning from Human Feedback, where the learner seeks a policy with a large worst-case win rate against both an adversarial competitor and any preference kernel lying in an ambiguity set around a nominal preference. When the ambiguity set captures the uncertainty in preferences, the resulting hard-constrained robust objective directly yields a certified lower bound on worst-case performance. However, we note this problem is computationally challenging to optimize, and to address this, we introduce a four-player primal-dual proxy game involving the leader policy, follower policy, adversarial kernel, and dual variable, and develop a single-loop optimistic mirror descent-ascent algorithm for this game. We show that the proxy always lower-bounds the truncated hard-constrained objective, quantify the proxy-to-hard gap, and characterize an exactness condition under which the proxy recovers the robust objective. We then prove an $O(1/\sqrt{T})$ convergence rate for the proxy-game duality gap, which implies a near-optimal robust policy for the original robust objective. Experiments on controlled tabular games and LLM alignment with uncertain preference further validate the convergence theory and show improved performance over nominal baselines.
Chat is not available.
Successful Page Load