Principled Top-$k$ Selection for Language Models with Hybrid Gradients
Xuchen Gong ⋅ Junfei Sun ⋅ Tian Li
Abstract
Selecting the best $k$ items out of $m$ candidates is a critical component of modern large language model systems, such as document selection in Retrieval-Augmented Generation (RAG) and expert routing in Mixture-of-Experts (MoEs). However, training these selection modules remains challenging due to weak gradient signals and suboptimal exploration/exploitation tradeoffs. Furthermore, prior works often rely on heuristics, lacking principled objectives and approaches that explicitly model and solve the top-$k$ selection problem. In this work, we propose a principled objective for training selection modules, whose gradient naturally provides richer training signals in a hybrid form---containing both a supervised component and policy gradients. We show that the selection problem becomes harder as $m$ increases, and our algorithm converges at rate $O(1/\sqrt{T})$, with the optimal upper bound achieved by balancing between bias and variance. Practically, we apply our method to a set of tasks involving top-$k$ selection, including synthetic regression tasks, RAG, and MoE systems, showing that our method outperforms strong baselines in language modeling perplexity and QA task accuracies.
Chat is not available.
Successful Page Load