Replicas-as-Variables: A Planner for Throughput-Optimal Routing in LLM Deployment
Abstract
Deploying Large Language Models on heterogeneous GPU clusters is often bottlenecked by rigid parallelism strategies that fail to fully exploit hardware asymmetry. To address this, we introduce ReVar, a holistic framework that optimizes replica placement and activation routing by treating replica counts as primary decision variables for all model components, including both dense stages and MoE experts. By modeling inference as a coupled resource-allocation and flow-routing problem, ReVar replaces static pipelines with fluid execution graphs that utilize uneven replication and multi-path routing. Supported by rigorous theoretical foundations, including proving the problem's NP-hardness, providing a MILP formulation for throughput upper bounds, and deriving closed-form solutions for constrained regimes, ReVar delivers significant practical gains. In evaluations, it achieved the highest throughput in 14 out of 15 settings, demonstrated up to a 3.1 times improvement in steady-state throughput in mixed-GPU environments, and was successfully validated on a real physical heterogeneous cluster.