PRICE: Adaptive LLM Test-time Compute under a Token Budget and Its Closed-form Pareto Frontier
Zhenyu Wang ⋅ zhu ⋅ Yifan Hu
Abstract
Test-time compute improves LLM accuracy by drawing more rollouts per query. A deployer pays tokens for these extra rollouts and must trade cost against accuracy. This raises a question: given a fixed token budget, how should it be spent to maximize accuracy? We formulate this as a token-level constrained optimization problem over policies that choose, per query, both the number of rollouts and the voting rule that turns these rollouts into an answer. Since the token budget determines both decisions, we call the optimal policy **PRICE** (**P**riced **R**ollouts and **I**nference-time voting-rule **C**hoice under a token budg**E**t). Theoretically, PRICE's cost--accuracy Pareto frontier dominates that of every fixed voting rule at every budget. We establish, to our knowledge, the first closed-form characterization of the Pareto frontier for the LLM test-time compute, identifying the attainable accuracy ceiling and the convergence rate at which the accuracy approaches this ceiling as the budget grows. Empirically, at matched accuracy on MATH-500, PRICE on the oracle setup spends ${\sim}6.8\times$ fewer tokens, while its practical version spends ${\sim}3.1\times$ fewer tokens than existing baselines.
Chat is not available.
Successful Page Load