Diagnosing and Mitigating Evidence Narrowing in Graph-Based Reranking
Min Zhang
Abstract
Graph-based reranking has been explored for improving retrieval-augmented generation (RAG) by propagating relevance over document--document similarity graphs. While graph-based propagation can improve ranking, it may also concentrate retrieved results within a small number of semantically similar regions. We refer to this phenomenon as evidence narrowing: a reranker may preserve relevance while reducing the semantic breadth of the retrieved evidence. We study this effect on SciFact, NFCorpus, and TREC-COVID from BEIR. For each query, we fix the BM25 top-100 candidate set and apply all rerankers within this set. Candidates are represented using normalized sentence-transformer embeddings and clustered with KMeans using approximately $\sqrt{N}$ clusters. We evaluate intra-list similarity (ILS), semantic cluster coverage (ClustCov), largest-cluster share (LCS), and nDCG@10. To diagnose the role of graph structure, we compare SemanticPPR with random-edge, degree-preserving rewired, and edge-weight-shuffled PPR controls while holding candidates, BM25 scores, and PPR parameters fixed. Across all three datasets, SemanticPPR increases ILS@10 by 0.096--0.112 and decreases ClustCov@10 by 0.074--0.102 relative to BM25. On TREC-COVID, it improves nDCG@10 by +0.080 while increasing ILS by +0.112 and reducing cluster coverage by -0.074, showing that relevance gains can coexist with narrower retrieved evidence. Graph-control comparisons consistently associate stronger narrowing with semantic graph topology rather than generic diffusion or degree structure alone. Diversity-aware PPR partially mitigates this effect, revealing a relevance--diversity trade-off. These results motivate evaluating graph-based RAG rerankers jointly for relevance, evidence coverage, and concentration.
Chat is not available.
Successful Page Load