Equip, Don't Replace: Query-Time Memory Organisation with Retrieval as a Tool
Abstract
A personal AI assistant must answer questions whose evidence lies scattered across years of its user's photographs, videos, and emails. On ATM-Bench, the first benchmark for this setting, state-of-the-art memory systems achieve an accuracy of under 20% on such multi-evidence questions. We show the bottleneck is not retrieval alone. Our hybrid, query-adaptive retriever sets a new state of the art on both benchmark splits. LLM agents that replace it fail at search itself. Even oracle retrieval leaves most multi-evidence questions unanswered, indicating that answer generation falls short as well. We therefore equip rather than replace. REMMI is an agent that organises relevant memory into events at question time, calling the optimised retriever as a tool alongside its own search. REMMI raises multi-evidence accuracy from 32.8% to 49.8% with a strong agent model, while reducing tokens per point of accuracy by roughly 30%. Our results yield a placement principle: retrieval pipelines are suited to finding single pieces of evidence, agents to composing many, and tool use connects the two.