When Should I Pay for Better LLMs? Stochastic Optimization with Switching Mixed-Fidelity Oracles
Abstract
Large language models (LLMs) are increasingly used upstream of decision-making to clean data, detect label errors, and reconstruct missing values. Yet LLM-assisted data curation creates a new operational tradeoff: higher-fidelity models and richer prompts can be substantially more expensive, while the value of this additional fidelity depends on the downstream optimization accuracy. This raises the central question of this work: \emph{when should an optimizer pay for a better LLM?} We model each model--prompt pipeline as a stochastic first-order oracle with its own cost, state-dependent variance, and bias, and study how to switch among such oracles as optimization progresses. We propose a restarted, cost-aware policy that adaptively selects oracle fidelity across optimization stages. Its cumulative cost is no larger than that of the best fixed oracle in hindsight, and can be substantially smaller. When oracle biases are unknown, we develop a bias-aware scheme that uses a trusted gradient only to certify progress in optimization. Synthetic and LLM-completion experiments exhibit the predicted switching behavior: the adaptive policy skips several expensive LLM oracles and substantially reduces the cumulative cost relative to fixed high-fidelity alternatives, while achieving comparable accuracy.