Evaluating Memory Safety as a Trajectory: A Trigger–Probe Protocol
Abstract
Persistent memory carries information across tasks, so an agent's safety can change as its history grows. We therefore view memory safety as a trajectory rather than a property of a single interaction. Measuring this trajectory is difficult because cumulative violations can rise as the input stream changes, even for a stateless agent with memory disabled. We introduce a trigger--probe protocol that keeps test queries fixed while memory grows. At each checkpoint, we rebuild memory from a longer history, evaluate the same benign probes without storing them, and compare each response with a matched no memory baseline, which we call NullMemory. Across seven memory architectures on synthetic and real correspondence streams, memory-induced violation rates rise with exposure for most architectures under benign use. The pattern remains under block and full shuffling of the history, suggesting that accumulated content contributes beyond encounter order. Our protocol provides a controlled way to evaluate memory safety as a trajectory over long-term use.