Scaling Offline Inference Compute Improves Online Inference Efficiency
Shayan Talaei ⋅ Jirat Chiaranaipanich ⋅ Amirreza Zeinali ⋅ Azalia Mirhoseini ⋅ Amin Saberi
Abstract
Can inference compute be performed before a query arrives and reused to reduce the compute required afterward? We study this question in settings where task-relevant context, such as a codebase, database, example set, or partial problem, is available before the exact query is known. A query-blind model spends an offline inference budget on this context and transfers its full reasoning trace to a solver that subsequently receives the query and an online inference budget. Across nine benchmark panels spanning mathematics, algorithmic reasoning, software engineering, and SQL, we jointly scale offline and online inference compute. Offline inference improves accuracy most when online compute is scarce — by up to $+0.69$ on AIME 2024 at the smallest online budget — increases token efficiency, and shifts tool use out of the latency-critical online phase. A linear effective-compute model, $T_{\mathrm{eff}}=T+k_bR$, collapses all nine benchmarks onto a single sigmoid ($R^2=0.945$), while online headroom and context--query relatedness predict where offline inference pays (AUC up to $0.98$). These results suggest that resource-aware inference can be viewed as a temporal compute-allocation problem, in which reusable pre-query computation can substitute for per-query computation.
Chat is not available.
Successful Page Load