Target-Budget Best-Function Identification for Scaling-Law-Guided Model Selection
Iurii Dorn
Abstract
Scaling laws are increasingly used to pick which expensive training run to scale up. The practical question — which candidate is best at a fixed future compute budget $C^\star$ — fits neither classical best-arm identification nor minimum-oriented functional bandits. We formulate it as the *target-budget best-function identification* (TB-BFI) problem, with two natural flavors: TB-BFI-FC asks for a $\delta$-correct certificate of the winner at $C^\star$, and TB-BFI-FB asks for low target-budget simple regret under a hard exploration budget. For the shared-exponent power-law class we prove uniform target-budget validity, a finite-time stopping bound for the cost-aware TB-LUCB rule, and a per-arm-exponent perturbation extension. We then close two structural gaps. **(a)** We extend the theory *beyond* shared-exponent to the joint $(N,D)$ Chinchilla law with shared but *unknown* exponents, via a two-stage protocol with explicit pilot sample complexity and a quadratic-in-$\epsilon$ centred-design perturbation theorem. **(b)** We give the first positive TB-BFI-FB result: a cost-aware Sequential-Halving algorithm whose cost-weighted simple-regret upper bound matches a new information-theoretic lower bound up to the standard $K \log K$ adaptive-vs-oracle gap. Empirically, theorem-aligned synthetic experiments certify the $C^\star$-winner in $95$–$100\%$ of trials; on a char-level transformer learning-rate-crossover our extrapolator picks the right arm in $3/3$ seeds where Successive Halving picks $0/3$; and on a flip-rank Chinchilla instance the joint-law procedure reaches $70$–$87\%$ accuracy where the univariate $c = N \cdot D$ baseline scores $0\%$.
Chat is not available.
Successful Page Load