Learning Complementary Modes for Pareto Set Optimization
Abstract
The question of whether new optimizers can outperform Adam is usually framed in terms of better parameter updates. We study a complementary regime in which the main bottleneck arises before the update rule: the task provides only a non-differentiable, non-additive set-level utility, making it difficult to construct informative learning signals for individual components. We propose a two-stage framework that augments Adam with RL-based marginal credit assignment. The first stage learns a shared library of complementary latent modes by rewarding each mode according to its marginal contribution to the output set; the second trains an amortized, budget-conditioned selector using marginal utilities derived from the same objective. We instantiate the framework for finite-budget Pareto-front approximation in multi-objective shortest-path problems using dominated hypervolume and a frozen Qwen3-4B backbone. Across four synthetic settings and the public eight-objective GridK8 benchmark, our method achieves higher normalized hypervolume than all evaluated baselines for every constrained budget, with the largest gains under tight budgets. These results suggest that advancing optimization may require not only improving how gradients are used, but also improving how optimization signals are constructed for structured, black-box objectives.