Obligation-Aware Long-Term Agent Memory: Scoped Operations and Failure Localization
Abstract
Long-term agent memory turns conversations into persistent records that can change after the interaction ends. In regulated settings, record-level obligations govern both admission and later handling. Existing evaluations assume a fixed retrieval or deletion target, so they cannot represent several obligations over different records or locate failures between a record-level decision and persistent state. We introduce nine composable operational primitives, each defined over an explicit record scope, and a three-stage model that separates decision errors from failures of system capability or execution. We evaluate this approach across memory systems with synthetic financial-services scenarios based on public US and EU legal and regulatory sources. Required outcomes often depend on application controls that direct calls can bypass. Even with the expected operations supplied, records can disappear from retrieval while remaining in the backing store or history. An exploratory probe also finds that the model often adds primitives beyond the required set. Obligation-aware memory therefore requires both record-scoped decisions and controls that reliably enforce them as memory changes.