Adaptive Sequential Retrieval in RAG via Generative Modeling and Reinforcement Learning
Abstract
Retrieval-Augmented Generation (RAG) has emerged as a powerful and widely adopted paradigm for grounding the responses of Large Language Models (LLMs). By retrieving relevant context from an external knowledge source, RAG enables LLMs to generate responses based on external information without requiring additional fine-tuning. However, effective RAG requires not only retrieving sufficient information for accurate response generation, but also avoiding unnecessary and redundant context that increases input token consumption and computational cost. This makes retrieval an important resource-aware decision-making problem: the retriever must acquire sufficient evidence while minimizing the amount of information passed to the generator. Existing retrieval approaches are largely task-agnostic and static, relying on similarity-based measures such as cosine similarity to rank and retrieve a fixed set of relevant chunks. However, the relationship between a query and the set of supporting chunks can be task-dependent and significantly more complex than can be captured by static similarity measures. Moreover, static retrieval provides limited flexibility in deciding when sufficient evidence has been acquired and retrieval should terminate. To address these limitations, we first propose a generative formulation of the RAG process and subsequently model retrieval as a Markov Decision Process (MDP). We employ a Deep Q-Network (DQN) to learn an adaptive retrieval policy that sequentially selects chunks and determines when to stop retrieval. The resulting formulation explicitly captures the trade-off between response quality and retrieval cost, encouraging the agent to acquire sufficient evidence while avoiding redundant chunks. Experimental results demonstrate that the proposed approach consistently outperforms conventional static retrieval strategies in terms of success rate and evidence coverage, while adaptively controlling the number of retrieved chunks.