MARCS: Marginal-Cost Routing and Adaptive Priority Control for LLM Inference
Abstract
Large language model (LLM) serving couples two decisions: route each request to a serving replica, then order that replica’s local queue into GPU batches. Continuous batching and chunked prefill link these decisions through token workload, key-value (KV) cache occupancy, and incumbent latency. We propose MARCS, pairing marginal-cost routing with Thompson Sampling over first-in-first-out (FIFO) and aging shortest-job-first rules. We establish a finite-time decomposition of its additive mean-latency surrogate excess over a matched two-level oracle into routing-score error, softmin randomization, priority regret, and trajectory coupling. On real GPUs, MARCS outperforms every baseline on the mean–tail tradeoff; calibrated simulations trace the operating curve, evaluate decode-only routing, and quantify routing–priority composition.