Can We Answer More with Less Context? Harness Engineering for Agentic RAG with Recursive Language Models
Abstract
Retrieval-augmented generation faces a scaling paradox: retrieving more evidence improves recall, yet answer quality plateaus once the accumulated context dilutes the generator's attention, leaving a retrieval-generation gap. We address it by conducting harness engineering around the LLM, especially applying Recursive Language Models (RLMs) at multiple stages of the agentic RAG pipeline so that raw text is distilled before it reaches the root LLM. We instantiate this as RLM-RAG, with three components. An RLM context reader holds retrieved candidates as objects in a code environment instead of prompt text, so the root LLM's context stays small however much is retrieved and exact operations over graphs, tables, and code run as code. Hive-mind retrieval agents turn that reader into an agent: a master RLM calls the retriever from its own program, hands each candidate to a sub-agent that returns a short summary, and writes the answer itself, so the agent that plans retrieval never reads a raw passage. Belief-propagation context assembly then reorders what the agent collected, scoring each piece of evidence by how many of the agent's own queries corroborate it. We evaluate on BrowseComp-Plus (a long-context multi-hop benchmark with large documents) and STaRK-Prime (a semi-structured benchmark containing both documents and a large knowledge graph), ablating harness variants and demonstrating their effectiveness. We also find that no single configuration wins everywhere: which stages pay off depends on the query, the data, and the task, and our ablations give concrete guidance for choosing among them.