Reward-Proportional Policy Optimization for Diversity-Driven Search
Abstract
Diversity-driven search seeks to maintain coverage over multiple high-quality outcomes rather than concentrate on a single optimum. Standard expected-return reinforcement learning often collapses onto a narrow subset of rewarding outcomes, while existing Inverse Probability Scaling (IPS) relies on within-group terminal frequencies that become uninformative when repeated outcomes are rare. We introduce Multiplicity-Aware Inverse Probability Scaling (MIPS), which estimates terminal probabilities by combining trajectory likelihoods with a learned multiplicity-allocation model over trajectories leading to the same outcome. MIPS integrates directly with Group Relative Policy Optimization (GRPO) without requiring specialized flow-matching objectives. Across hypergrid, molecular synthesis, and phylogenetic tree construction, MIPS-GRPO preserves substantially broader outcome coverage than GRPO and IPS-GRPO while closely matching reward-proportional targets and remaining competitive with GFlowNet-based methods. More broadly, this provides a foundation for diversity-preserving RL in post-training and robustness or safety applications.