Catching RLHF Failures before Training
Matteo Lippi ⋅ Santiago Tomas Aranguri Diaz
Abstract
RLHF plays a key role in post-training, but the most reliable way to debug an RLHF setup, including its reward model (RM), is still running the training end to end: static RM benchmarks and Best-of-N (BoN) sampling correlate poorly with downstream outcomes. We introduce Lookahead, a sampling strategy that anticipates training failures before they occur, using only the base model and the RM. At selected decoding positions, each candidate token's logits are shifted by the reward that a rollout continuing from it obtains, with a coefficient $\alpha$ controlling the optimisation pressure applied. At high $\alpha$, the RM's preferences are amplified until its flaws become visible. On a Tulu 2.5 setup, Lookahead outperforms BoN at predicting which IFEval prompts degrade after training ($\mathrm{F1}^{-}$ 0.667 vs.\ 0.498, a gap no BoN sample budget up to $N = 2000$ closes), and correctly flags every instruction family that training degrades, on each of which BoN predicts improvement. Our results show that much of what RLHF will break is already determined by the (base model, RM) pair, and can be read out before training begins.
Chat is not available.
Successful Page Load