RAG in a Trenchcoat: When Minimal Memory Is Enough for Agentic Systems, and When It Isn’t
Jingyu Liu ⋅ Zongze Li ⋅ Zach Xu ⋅ Zhanhui Zhou ⋅ Tahseen Rabbani ⋅ Dawn Song ⋅ Ce Zhang
Abstract
Agentic memory research has accumulated extraction, supersession, reflection, and graph synthesis on top of retrieval-augmented generation. Yet on the field's three established benchmarks, almost none of those primitives earn their cost: We show this with Bridge, a deliberately minimal RAG memory with chained observation notes, hybrid BM25 + cosine retrieval, and chronological rendering, which reaches parity within bootstrap CI with the prior published SOTA on LongMemEval-S and BEAM, and beats the best memory baseline on EverMemBench by $+5.7-6.3$pp at matched reader. It runs at $\sim5$-$7$K reader-prompt tokens per query and a single LLM call for each ingestion or query. We argue this parity is itself the result: standard benchmarks score what memory retrieves, not what memory enables when the agent acts. To test the latter, we introduce five adversarial probes (Rule Application $\pm$, Pattern Induction style / procedural, Dynamic Multi-turn) judged at generation time by a cross-vendor verification chain; each can only be solved by applying a remembered rule or pattern the probe never mentions. Across 5 probes $\times$ 20 scenarios, no architecture wins universally: atomic-fact extractors win rule application (LangMem leads Bridge by $+45$pp on RA$^+$, with non-overlapping bootstrap CIs), while Bridge's block-summary memory wins pattern induction (leads the strongest baseline by $+33$pp on PI-P). The right memory architecture is mechanism-dependent, and the field's evaluation paradigm fails on both ends: standard benchmarks fail to distinguish a minimal memory from a complex one, and application-grade probes that do distinguish find no universal winner.
Chat is not available.
Successful Page Load