Adaptive Pairwise-Feedback Sampling for Multi-Prompt KL-Regularized Policies
Abstract
Pairwise feedback is often available when calibrated rewards, a reliable reward model, or repeated policy retraining are unavailable or too costly. We study how to improve response policies for a collection of prompts when the only supervision is a stream of noisy comparisons between candidate responses. Rather than first converting comparisons into a learned reward model, our method uses them directly to refine a response policy while staying close to a trusted reference policy. The central challenge is that a finite feedback budget must be shared across prompts: some prompts are more important in deployment, while others require more feedback before their response policies become reliable. We derive an ideal allocation of comparisons and develop an online procedure that learns this allocation from the feedback it receives. We prove that the policy returned by the method improves at the same leading rate as if the prompt-specific difficulties were known in advance. The result provides a direct, theoretically grounded alternative for using fresh pairwise feedback when the candidate responses are finite and already available.