XpertGPT: Efficient Language Modeling via Multi-Scale Sparse Expert Routing
Abstract
We introduce XpertGPT, a sparse decoder-only Transformer that routes tokens through parallel expert pathways operating at multiple contextual scales (sliding-window spans from 4 to 64 tokens), alongside a global-attention stream, using Expert-Choice routing for load-balanced token assignment. Under a data-constrained training regime, XpertGPT activates only 32.1M of its 52.6M parameters per token-a 67\% reduction relative to a comparable 98.4M-parameter standard baseline-and requires roughly one-third the theoretical FLOPs per forward pass, while reaching comparable downstream quality across grammatical, world-knowledge, and reasoning benchmarks. Controlled ablations isolate the contributions of individual architectural components, showing that global context and multi-scale expert pathways contribute to the observed performance differences.