JMAP: Joint Multi-Document Attention with Clustering-Based Adaptive Context Pruning for Retrieval-Augmented Generation
Jasurbek Ibragimov ⋅ Ahmet Aksoy
Abstract
The context returned by a retrieval-augmented generation (RAG) pipeline often contains irrelevant and distracting documents. This noisy context increases the latency and computational cost of the reader LLM and degrades answer quality when the model is instructed to rely on the provided text. Rerankers operate on whole passages and therefore cannot remove irrelevant content at the granularity of individual sentences, while existing sentence-level pruners score each passage or document in isolation and thus fail to capture evidence that is distributed across several documents. We present JMAP, a 149M-parameter ModernBERT-base cross-encoder that scores all sentences of a multi-document context jointly in a single 8,192-token forward pass and dynamically selects the cut-off threshold by clustering the scores. On HotpotQA, 2WikiMultihopQA, MuSiQue, and TriviaQA, JMAP retains the evidence sentences more accurately than every baseline, and its pruned context improves the answer F1 of a Qwen3-8B reader over the unpruned context on all four datasets while removing 81 to 87% of the tokens. The pruner requires only 1.6 to 2.0% of the reader's full-context compute and reduces end-to-end latency by 1.24 to 1.45$\times$.
Chat is not available.
Successful Page Load