Metered or Provisioned? Joint Routing and Capacity Planning for Resource-Aware Inference
Abstract
Serving stacks increasingly span heterogeneous inference resources: specialized models on provisioned capacity that is cheap per request but expensive to keep ready, and hosted foundation models that need no warm-up but bill per token. Deciding which resource a request deserves has to weigh quality against how that resource is billed. We study a hybrid resource-aware serving design (HRS) that jointly routes requests and sizes capacity between provisioned tenant-specific DNNs and a metered hosted LLM. A cheap semantic score estimates the metered share. That share is carved out of forecast load before the warm fleet is sized, and the LLM absorbs cold-start and overflow. Our case study is a multi-tenant enterprise recommender whose per-tenant artifacts take about 50 minutes to materialize, which forces a large always-warm fleet. We measure both engines on the live fleet and replay routing and capacity on production-derived traces. Offline replay of HRS reaches traffic-weighted Hit@5 0.902 against 0.913 for the always-warm fleet, at USD 13.2K against that fleet’s profiled USD 65.8K, with a 34-pod peak against 122 pods held.