Event Structure and Memory in LLM Agent Trajectories
Abstract
LLM agents notoriously drop earlier constraints as multi-task sessions grow. Human cognition offers a candidate mechanism: experience is segmented into events, and event boundaries reshape memory. We test whether this account transfers to agents by pairing a controlled multi-task tool-use environment (MultiDesk) with an audit-grade analysis of residual-stream trajectories from a 7B agent. Three findings. (1) Task-frame switches are representational events: at turn resolution, the state displacement at a switch exceeds matched-turn nulls by about five standard deviations, replicated across two independently generated corpora, and turn states cohere within frames. (2) Boundaries are not the behavioral cost: without competing facts, access to a planted fact is flat across matched distances; with matched interference, access falls log-linearly with token distance, while the per-boundary cost is positive but statistically unresolved. (3) Validity is fragile: a degenerate regression manufactured a spurious distance law, fact decoding collapsed from 0.69 to 0.20 once environment content was made independent of the target fact, and a fixed segmentation cap silently corrupted boundary alignment on long trajectories. We argue that claims about agent trajectories require deconfounded designs and validity gates throughout.