BATCHPACE - Global Budget LLM Routing under Batched Arrivals
Abstract
We study LLM routing when queries arrive in fully observed batches, decisions are irreversible across batches, and one escalation budget is shared over the horizon. We propose \BP, which combines predicted query-level marginal gains with a Bellman continuation value to decide both which queries to escalate and how much budget to spend in each batch. Under an oracle model, the policy is Bayes optimal, admits a dynamic shadow-price rule, and benefits monotonically from larger information windows; an end-to-end bound separates prediction error from continuation-value error. Stationary synthetic and MMLU-Pro-derived experiments show that the benefit depends on exploitable within-batch heterogeneity and is strongest when high-value opportunities cluster within observable batches.