The Statistical Illusion of Rejection Sampling in LLMs: Bridging the Gap Between Heuristic Truncation and True Alignment
Abstract
Rejection Sampling Fine-Tuning (RSFT) is widely used in LLM alignment, yet the procedure practitioners call “rejection sampling” is not rejection sampling in any statistical sense. It is a rank-and-truncate heuristic that selects the highest-scoring candidate from a batch and discards the rest. We prove that this induces a distribution over fine-tuning data that converges to a point mass on the reward argmax as the candidate pool grows, with reverse KL divergence from the theoretically optimal KL-regularized policy diverging to infinity—a formal account of the mode collapse observed in practice. To correct this, we derive Distribution-Preserving Rejection Sampling (DP-RS), which applies the von Neumann envelope construction to the RLHF objective. Setting the envelope at the empirical batch maximum causes the intractable partition function to cancel, reducing the problem to a single probabilistic acceptance step per candidate. We prove that the practical algorithm converges in total variation to the Boltzmann-optimal policy as batch size grows, and establish the KL-reward trade-off controlled by the temperature parameter β. Experiments on Anthropic HH-RLHF, AlpacaFarm, and GSMK show that DP-RS achieves higher win rates than hard RSFT while maintaining generation diversity close to the reference policy.