Trust Momentum Gradient for Teaching Trajectories Optimization
Xiaofeng Cao ⋅ Keyu Gu ⋅ Pengkun Wang ⋅ Jingcai Guo ⋅ Jielong Yang ⋅ Xiangtao Li ⋅ Jiangchao Yao ⋅ Wei Ye
Abstract
Teaching a student model along non-convex trajectories often leads to convergence collapse, resulting in spurious convergence behavior, particularly under SGD. In this paper, we propose TrustMomentum, a gradient solver that restricts loss optimization to locally trusted momentum gradients, avoiding potential cumulative deviation in near-descent iterative optimizations. Theoretical analysis guarantees an $\mathcal{O}(1/\sqrt{T})$ convergence rate, where $T$ denotes the number of iterations, yielding consistent sample complexity as SGD. Extensive experiments including machine teaching, knowledge distillation and LLM fine-tuning confirm smoother convergence with superior stability while preserving momentum acceleration.
Chat is not available.
Successful Page Load