Risk-Aware Best-of-$N$
Yichun Hu ⋅ Meng Qi
Abstract
Best-of-$N$ (BoN) is a widely used inference-time alignment method that samples $N$ candidate responses from a fixed reference policy and returns the response with the highest predicted reward, where the reward model is typically trained from pairwise preferences by maximizing a Bradley-Terry (BT) likelihood. The performance of the BoN policy depends on the downstream criterion used to evaluate the selected response. In this work, we consider a broad class of deployment criteria defined on the BoN reward-gap distribution and investigate if introducing risk-aware weights into BT training would improve standard BT. We first characterize the asymptotic performance of the weighted BT models and demonstrate the bias-variance trade-off caused by risk-aware weighting. As a corollary, we identify when risk-aware weighting improves upon standard BT across mild, balanced, and severe model misspecification regimes. Motivated by these theoretical findings, we then develop a data-driven procedure that uses preference labels and reference-policy simulation to construct an objective-specific local weighting direction. We evaluate our procedure in a semi-synthetic experiment based on the five-response llama3-ultrafeedback-armorm snapshot and find the predicted bias--variance tradeoff: weighting incurs an efficiency cost under correct specification but becomes increasingly beneficial as misspecification grows.
Chat is not available.
Successful Page Load