What Survives State Reset? Controlled Evidence for Replay and Structured Long-Term Memory
Abstract
Long-context agents can retain information in transient state, replayed experiences, or a persistent memory module, but a recall score alone does not identify which carrier survived. We introduce a reset-survival protocol that clears short-term state before held-out queries and audits replay, routing, and decoy rejection separately. On a frozen sequence model, a bounded structured long-term memory (LTM) written during an offline Sleep phase reaches 1.0000 across three fixed seeds, versus 0.7500 without cross-session replay, 0.5000 without Sleep, and 0.2865 for a continued low-rank adapter. Shuffled learned signatures fall to 0.5339 and random keys to 0.0738. A direct attempt to consolidate the same traces into weights fails the joint value and decoy criterion. A separate integration probe with a frozen distilgpt2 tokenizer at 16K and 32K inputs reaches 0.8125 on a 25-token candidate task, tied with a RAG-like control. The evidence supports replay plus a separate structured readout on this controlled surface; it does not establish open-ended language-model memory, retrieval superiority, or general long-context reasoning.