Revisiting Gradient-Normalized Linear Scalarization for Multi-Task Learning
Abstract
Multi-task learning requires optimizing multiple objectives whose gradient magnitudes can differ substantially, allowing standard linear scalarization to be dominated by tasks with larger gradients. We revisit a simple alternative, termed Gradient-Normalized Linear Scalarization (GNLS), which normalizes each task gradient before averaging the resulting directions. GNLS is invariant to task-specific gradient scales, requires no auxiliary optimization subproblems, and can be readily combined with standard gradient-based optimizers. Across widely used multi-task benchmarks, we observe a recurring aggregate directional-consistency phenomenon: despite the presence of pairwise gradient conflicts, the average normalized task gradient typically maintains a positive projection onto the gradient of the equally weighted objective throughout training. Building on this empirical property, we establish asymptotic convergence guarantees for GNLS equipped with the Adam-type optimizer on smooth nonconvex objectives. Experiments on Cityscapes, NYU-v2, QM9, and CelebA, spanning 2 to 40 tasks, show that GNLS consistently outperforms standard linear scalarization and achieves competitive or superior task-balanced performance compared with more complex multi-task optimization methods.