LOSCAR-SGD: Local SGD with Communication-Computation Overlap and Delay-Corrected Sparse Model Averaging
Abstract
Communication is a major bottleneck in distributed learning. Three standard remedies are compression, local training, and communication--computation overlap; methods that combine them are used in practice, but little theory covers all three at once. We study a heterogeneous-compute setting in which workers may take different numbers of local steps, and analyze LOSCAR-SGD, a Local-SGD method that communicates only a sparse subset of model coordinates and keeps optimizing while communication is in flight, merging the delayed sparse average through a delay-corrected rule that preserves the progress made during the overlap phase. For smooth non-convex objectives we give convergence and wall-clock time guarantees showing how sparsity, overlap, and worker heterogeneity enter the rate, with the price of overlap confined to higher-order terms. To the best of our knowledge, this is the first analysis of this combination. Experiments show that overlap consistently reduces simulated training time, while the preferred merge rule is regime-dependent.