Hierarchical Soft Preference Learning for Pluralistic Alignment
Motahareh Sohrabi ⋅ Seyed Meraj Hashemizadehaghda ⋅ Dylan Hadfield-Menell
Abstract
Soft Preference Learning (SPL) (Slocum et al., 2025) aims to represent heterogeneous preferences by decoupling the entropy of the policy from its cross-entropy against a reference policy. We show that, applied to autoregressive language models, the resulting stationary distribution need not be normalizable over the space of sequences. We attribute this failure to a mismatch between the level at which SPL encourages diversity, individual response strings, and the alternatives over which proportional representation is defined. Based on this observation, we propose hierarchical SPL (H-SPL), which applies SPL across a known finite partition of alternatives and retains KL regularization within each alternative. H-SPL has a proper population optimum for every $\beta\ge0$ and admits a DPO-style training objective. In a controlled three-style experiment, H-SPL at the theoretically proportional setting $(\alpha,\beta)=(1,0)$ achieves lower total variation distance to the target than tuned DPO and SPL across uniform and skewed distributions.
Chat is not available.
Successful Page Load