GridFlow: Decentralized Block-Based Flow Matching
Abstract
The infrastructure required to train large generative models is increasingly concentrated in a small number of hyperscale data centers, excluding the vast majority of globally distributed compute. Meanwhile, communication-free decentralized training has been demonstrated along two distinct axes: data, through independent experts trained on separate shards, and time, through independent blocks trained on different segments of the generative trajectory. Alone, each axis limits either the degree of decentralization or the efficiency of inference. We instead show that data and time form independent axes of the same flow matching objective and can therefore be partitioned simultaneously, yielding a grid of blocks trained in complete isolation that multiplies the degree of decentralization while reducing per-device memory and inference cost. At matched training FLOPs, our best grids, which concentrate experts in the early and noisy stages of generation, outperform the monolith on CIFAR-10 and approach it on ImageNet-256, indicating a path towards large-scale generative modeling on smaller, distributed, high-latency hardware, such as through collaborative international efforts.