Rescaled Asynchronous SGD: Distributed Optimization under Data and System Heterogeneity
Ammar Mahran ⋅ Artavazd Maranjyan ⋅ Peter Richtarik
Abstract
Asynchronous stochastic gradient descent (ASGD) updates the model whenever a gradient arrives, so no worker ever idles. Under heterogeneous local data and computation times this may cause objective inconsistency: faster workers deliver more updates and the trajectory drifts toward their local objectives. We show that scaling workers stepsizes $\gamma_i$ to be proportional to their computation times $\tau_i$ restores the correct objective. We prove convergence to a stationary point of the true objective by analysing a continuous-time shadow trajectory, and derive a worst-case wall-clock time complexity. The leading term of this matches the theoretical optimum for parallel stochastic first-order methods, and scales with the arithmetic mean of computation times. Unlike existing methods, our proposed method does not introduce any memory overhead, gathering phases, buffers, or synchronization.
Chat is not available.
Successful Page Load