Test Time Search Requires Training for Diversity
Ryan Bahlous-Boldi ⋅ Isha Puri ⋅ Idan Shenfeld ⋅ Akarsh Kumar ⋅ Mehul Damani ⋅ Sebastian Risi ⋅ Omar Khattab ⋅ Zhang-Wei Hong ⋅ Pulkit Agrawal
Abstract
Language models are increasingly deployed alongside test-time search. We argue that we should therefore shift the role of RL post-training from converging on a single best response to producing a diverse pool of competent candidates. Standard policy gradient methods such as GRPO optimize a fixed scalar reward and drive the policy toward near-duplicate responses, erasing the diversity search needs. We propose Vector Policy Optimization (VPO), which exploits the fact that practical rewards are vector-valued, e.g., per-test-case correctness, per-criterion ratings, or per-hop credit. Instead of collapsing these into one scalarization, VPO samples randomized scalarizations from a distribution over the reward simplex, incentivizing candidates to specialize to different trade-offs along the Pareto front. The model emits multiple candidates per prompt, with each conditioning on the previous ones, allowing for directed diversification. VPO is a drop-in replacement for the GRPO advantage estimator. Across four domains, VPO consistently improves test-time best@$k$ over scalar baselines, with gains widening as the test-time budget grows.
Chat is not available.
Successful Page Load