The Geometry of Agent Memory: Forecasting Retrieval-Bound Memory Failures, and Characterizing Its Limits
Abstract
An LLM agent stores memories and retrieves them across sessions. We study memory systems that index and retrieve by dense-embedding similarity (extraction-then-retrieval and chunk-RAG designs), not symbolic, graph, or parametric memory. Such systems are usually compared by aggregate end-to-end QA accuracy, which does not say, before the system answers, which questions it will get wrong. A cheap pre-answer warning could let an agent abstain or defer the riskiest queries—and potentially broaden retrieval or fall back to the raw conversation—before producing an unsupported answer. We embed the query, the stored memories, and the conversation turns (which we add) in one space, and derive two feature families: query reachability (how close the query lies to the memory set) and memory–conversation correspondence (how well the memories cover the conversation). A linear classifier over these features forecasts per-question failure across five memory systems on LongMemEval-small. Performance is modest overall (AUROC 0.640) but strongly type-dependent: 0.841 for single-session-assistant questions, whose failures involve missing or unreachable memories, and near chance for reasoning-heavy categories, where the needed memories are usually already accessible. Reachability carries most of the signal (correspondence adds smaller, type-dependent gains), and an intervention shows query proximity affects the answer—the 15 closest memories beat random selection by about 7×. The split recurs on LoCoMo and is stable across four embedders. We also report a control that bounds the claim: once system identity is included alongside question difficulty, geometry adds no significant pooled signal (paired Δ = +0.002, n.s.), so its durable value is the within-type result and the mechanism, not a pooled predictor. Training and the offline scope analysis use correctness labels, but inference uses only pre-answer geometry and no extra generation call. Embedding geometry thus gives a cheap diagnostic for retrieval-bound failures and shows, offline, that reasoning-bound categories fall outside its scope.