TokenRouter: Efficient Serving System for Token-Level LLM Routing
Abstract
Large language model (LLM) routing is widely used to advance the cost-quality Pareto frontier of modern serving systems. While coarse-grained routing at the session or query level has been widely adopted in production systems, recent algorithm work shows substantial efficiency and quality advantages of finer token-level routing. However, efficiently serving token-level routed inference poses significant challenges to existing system design. Built on single-LLM assumptions, current systems struggle with step desynchronization and frequent batch admission delays under token-level routing, and they also impose high implementation complexity on developers. To address this, we propose TokenRouter, an efficient and user-friendly serving system for token-level routed LLM inference. TokenRouter shifts from the conventional request-centric paradigm to a model-centric one. Instead of treating each request as a synchronized decoding stream, TokenRouter launches a subserver for each candidate LLM and dispatches requests asynchronously according to routing decisions. Each subserver employs a delayed-batching scheduler, whose hyperparameters are derived mathematically from a throughput model of the system. Across diverse routing algorithms, workloads, and model pairs, TokenRouter achieves 1.83-64.29× higher decoding throughput than existing systems, substantially advancing the serving efficiency of token-level LLM routing.