The Renaissance of Line Search for Deep Learning?
Chongwei Chen ⋅ Yuxiang Wan ⋅ Wenjie Zhang ⋅ Sinian Zhang ⋅ Lingjie Su ⋅ Ju Sun
Abstract
The learning rate (LR) is a critical hyperparameter in deep neural network (DNN) training. In practice, predefined LR trajectories must be tuned through multiple runs of trial and error to fully realize the potential of DNNs, incurring substantial human effort and computation behind the scenes. In classical optimization, the LR is determined in a more principled way. One example is line search, which adaptively decides the LR at each step by checking whether it yields a sufficient decrease in the objective. Line search is impractical in DNN training due to its high computational cost. However, established DNN schedulers provide a useful structural prior: LR typically changes slowly piecewise. This allows us to perform line searches periodically rather than at every step, substantially reducing their computational cost. By combining periodic line search with two other standard DNN training heuristics---warmup and decay---we develop a novel LR scheduler that adapts robustly across tasks and reduces the need for manual LR-scheduler tuning. We conduct systematic experiments on LLM post-training and show that our line-search scheduler achieves performance within $5\%$ of the best baseline performance in most settings under minimum tuning budget.
Chat is not available.
Successful Page Load