Democratizing LLM Training Across Geo-Distributed Heterogeneous Compute: A Systems Position and Roadmap
Ziyue Luo ⋅ Jiaxuan Cai ⋅ Cedric Le Denmat ⋅ Srijith Nair ⋅ Fatemeh Nourzad ⋅ Rohith K Sudha ⋅ Qinhang Wu ⋅ Jifan Zhang ⋅ Sungjae Lee ⋅ Zhe Li ⋅ Peiwen Qiu ⋅ Rishabh Sharma ⋅ Sundararajan Srinivasan ⋅ Yinglun Xia ⋅ Xue Zheng ⋅ Zidong Liu ⋅ Bicheng Ying ⋅ Kaushik Chowdhury ⋅ Gauri Joshi ⋅ Yingbin Liang ⋅ Robert Nowak ⋅ Srinivasan Parthasarathy ⋅ Saurav Prakash ⋅ Balaraman Ravindran ⋅ Sanjay Shakkottai ⋅ Ness Shroff ⋅ Haibo Yang ⋅ Aylin Yener ⋅ Jia (Kevin) Liu
Abstract
The growing computational abilities of modern GPU-powered data centers have enabled industrial giants to train and deploy large language models (LLMs) on massive amounts of data. Yet the enormous scale of both models and training data makes centralized development increasingly difficult to sustain, while simultaneously excluding smaller institutions that possess relevant private data but lack equivalent infrastructure to participate in LLM training. In this paper, we organize our visions and positions on distributed LLM training around pretraining and fine-tuning paradigms, covering federated learning, split learning, and decentralized approaches, and assess their readiness for geo-distributed settings with heterogeneous hardware and unreliable connectivity. We then present a $\textit{unified}$ system architecture built around a cloud-hosted parameter server that coordinates heterogeneous clients over outbound gRPC connections, reusing the same communication, compression, and fault-tolerance infrastructure for both pretraining and a family of fine-tuning methods. We validate the design with real geo-distributed experiments on GPT-2 Medium and Llama3-1B pretraining across multiple university sites, showing stable convergence under realistic conditions. Finally, we identify open research directions spanning frontier model support, elastic fault tolerance, adaptive compression, federated alignment, privacy and trust, incentive design, and system-algorithm co-design that must be addressed to realize truly democratized LLM training.
Chat is not available.
Successful Page Load