CoRA-Bench: Evaluating Evidence Preservation by Context Selectors Under Token Budgets
Abstract
Context selection reduces the computational cost of long-input inference but can remove evidence required to answer a query. We introduce CoRA-Bench, a benchmark for evaluating evidence preservation under shared token budgets. We compare head and tail truncation, BM25, dense retrieval, dense retrieval with reranking, LLMLingua-2, and Provence across BABILong, LOONG Financial Spotlight, a custom NoLiMa diagnostic, and LongMemEval. Evidence-preservation results reveal task-dependent tradeoffs: BM25 achieves 0.914 supporting-fact recall on BABILong, dense retrieval achieves 0.898 needle recall on NoLiMa, and dense reranking achieves 0.990 answer-bearing-turn recall on LongMemEval. Complementary reader evaluations use overlap-aware packing, dataset-native answer metrics, and three runs on the same questions, with paired bootstrap confidence intervals. At a 32K budget, BM25 achieves 46.20% BABILong answer accuracy and dense retrieval achieves 71.13% LongMemEval accuracy. These results distinguish evidence preservation from downstream task success and characterize the conditions under which context selection supports long-context inference.