Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
Tianci Bu ⋅ Yuan Lyu ⋅ Zixi Chen ⋅ Chendong Song ⋅ Hong Liang ⋅ Tsepten Gurung ⋅ Yuwei Fan ⋅ Yinyu Ye ⋅ Zijie Zhou
Abstract
Data-parallel (DP) load balancing has emerged as a first-order bottleneck in large-scale LLM serving. When a model is sharded across devices via tensor parallelism (TP) or expert parallelism (EP) and replicated across many DP workers, every decode step ends in a synchronization barrier whose latency is set by the most heavily loaded worker; even modest persistent imbalance across DP workers compounds, step after step, into a substantial fraction of wasted compute. The problem is hard for reasons specific to LLM decoding: assignments are sticky (KV caches cannot be migrated), per-request loads grow over time, arrivals are non-stationary, and the router must decide within a sub-100ms decode budget over hundreds of waiting requests and tens of workers. We present BalanceRoute, a family of practical online routing algorithms that target this bottleneck. The first, BR-0, requires no prediction infrastructure and uses a piecewise-linear F-score that captures the sharp asymmetry between admissions that fill safe margin and those that overflow into the envelope; a two-stage decomposition keeps per-step cost compatible with millisecond-scale scheduling. The second, BR-H, generalizes BR-0 with a short, constant lookahead $H$ and a lightweight termination-classifier interface, extending the F-score to a horizon-discounted form. We deploy BalanceRoute on a 144-NPU cluster and evaluate against vLLM baselines on both a proprietary production trace and the public Azure-2024 trace. Across both workloads, BalanceRoute substantially reduces average DP imbalance and improves end-to-end serving throughput. An anonymized open-source release is available at https://anonymous.4open.science/r/BR-Family-on-vllm-acend-0F95/.
Chat is not available.
Successful Page Load