Adaptive Hyper-parameter Transfer for Adaptive Step Sizes
Abstract
To cope with the costs of tuning hyper-parameters to train increasingly large neural networks, hyper-parameter transfer suggests to perform the exhaustive search on a smaller, surrogate network and applying the best-performing hyper-parameters to train the large, target model. While the existing literature has been focusing on achieving zero-shot hyper-parameter transfer for a learning rate that is constant throughout training, in this paper we focus on the ambitious goal of transferring a whole sequence of adaptive step sizes from the training of the small network to that of the target one. We tackle this challenge by first studying the behavior of the step sizes yielded by monotone and nonmonotone line searches when training neural networks with different widths/depths. In general, we find out that there does not exist a fixed transfer parameter that allows to re-align their dynamics, least of all the value 1 (i.e., the zero-shot transfer parameter). We thus develop an adaptive transfer method that performs a line search on the small network, scales the resulting step size by an adaptive transfer parameter, and employs this value directly as an adaptive learning rate on the target network. By employing as an adaptive transfer parameter the ratio of the gradient norms of the two losses involved, we are able to show a first liminf convergence result for the deterministic version of our method. We test the proposed adaptive transfer algorithm against state-of-the-art stochastic line searches applied directly on the target model and show that the former is very competitive with the latter, especially in terms of per-epoch runtime.