Open, Collaborative, and Decentralized Training of Foundation Models
Abstract
Training large-scale foundation models today depends on massive, centralized GPU clusters that are inaccessible to most academic institutions, startups, and industries. This concentration of compute creates high barriers to entry, centralizes AI innovation, and limits broader progress in developing and studying frontier-scale foundation models. Open models are key to democratizing the know-how of frontier-model training. However, the cost of centralized training infrastructure largely excludes the broader research community from participating in open development at scale, leaving scaling efforts dependent on substantial centralized resources and offering the community limited opportunity to contribute to the training process itself. Decentralization and resource pooling provide a way to enable large-scale runs without this barrier by enabling model training across geographically distributed and heterogeneous devices, from coordinated inter-datacenter settings to consumer-devices connected over internet. This workshop will bring together researchers and practitioners to address the core technical challenges of this paradigm, including communication efficiency, asynchronous optimization, heterogeneous systems, fault tolerance, and security. Addressing these challenges enables large-scale foundation model training beyond the confines of a single datacenter, broadens participation in foundation model research, and provides complementary support for more scalable and collaborative open model development.