Train for Many, Update with One: Posterior Learning for Multi-Sample Reasoning
Abstract
Multi-sample inference improves reliability by generating K candidate solutions, increasing the chance that at least one succeeds within a limited inference budget. We formulate finite-K posterior learning, which trains a posterior over solver parameters for the collective success of a pool of K candidates. To optimize this objective efficiently, we propose Fenchel--Bregman (FB), which uses one posterior draw per update. First, an exact Fenchel reformulation turns the finite-K objective into a weighted one-draw update. Second, a closed-form Bregman update estimates the required weight from previous solver outcomes without additional posterior draws. We also analyze the conditions under which FB provides a more accurate update for finite-K posterior learning than the more expensive direct multi-draw training. Across challenging reasoning tasks, FB improves candidate coverage and achieves performance comparable to multi-draw training at lower training cost.