Memory as Evidence Composition
Abstract
Long-term agent memory is fundamentally an evidence-composition problem: under a limited context budget, a memory system must deliver not merely relevant items but the complete set of evidence the reader requires, and retrieval-augmented generation faces the same constraint. Most existing methods rank evidence items independently. This formulation has two limitations: (1) a highly ranked context can omit one fact required for a conjunctive or multi-hop answer, and (2) item-level metrics do not measure whether all required evidence reaches the reader. To address these limitations, we formulate retrieval and agent memory as budgeted evidence-set assembly and prove that assembly can attain the complete-evidence frontier at every budget precisely when the view family separates the store's evidence units. Quorum realizes this formulation: it constructs provenance-typed and query-derived views of the store, scores a shared candidate pool with a heterogeneous retrieval library and a learned candidate-level assembler, and selects evidence within the budget with a deterministic prefix packer. On HotpotQA, the largest reader gain occurs on the questions where Quorum completes an evidence set dense cosine leaves incomplete, raising exact match from 0.212 to 0.553; when both policies deliver complete evidence, reader accuracy is equivalent within the preregistered 0.02 margin. Across six benchmarks, Quorum improves complete-evidence retrieval over a dense-cosine baseline; HotpotQA joint@10 rises from 0.736 to 0.927. On LongMemEval-V2, Quorum raises accuracy over a standard RAG baseline from 0.385 to 0.530 while serving memory queries in 0.246 seconds on average, 46.9x faster than the nearest-accuracy comparator, and its margin over same-compression cosine is CI-positive at every tier of a 4x to 64x compression ladder (19.5 points at 32x, 19.1 at 64x).