Context Is Not Information: Common-Mode Suppression for Long-Horizon Agent Memory
Poorvi Patil
Abstract
Long-context foundation models are typically evaluated on whether they can find a relevant fact inside a very large context. We argue this understates the real failure mode of long-horizon agents: as trajectories grow, an increasing share of the context is redundant -- repeated observations, boilerplate tool output, restated reasoning -- and this redundancy is correlated across steps, not merely voluminous. We frame this as a signal-separation problem and propose to estimate and remove the dominant, common-mode subspace of an agent's trajectory before scoring steps for retention or retrieval, rather than relying on similarity, recency, or query-relevance alone. In a controlled synthetic setting where the number of decision-relevant observations is held fixed while trajectory length grows, common-mode suppression degrades more slowly than full-context, recency, similarity-to-centroid, query top-$k$, and MMR baselines, with the gap widening with length ($3$--$6\times$ the baselines by $L=2048$; these multiples reflect a synthetic setting favorable to the method -- real-trajectory evidence in the appendix is more conditional). A second experiment fixes length and sweeps a redundancy ratio $\rho$, isolating that the advantage tracks correlated redundancy, not context length per se. A distractor-type ablation -- matched so that only the correlation structure of the distractor differs, not its total variance or stationarity -- confirms this: against i.i.d. distractors, which have no shared subspace to remove, suppression gives no benefit over full-context, while against correlated distractors it gives a $1.7$--$2.7\times$ gain depending on how stationary the correlation is. We also find a crossover at short trajectories where naive scoring wins, and evaluate a causal (online) variant that captures much of the offline benefit at short-to-moderate length but fades at very large $L$. All quantitative results above are synthetic. Four real-trajectory evaluations of increasing scale confirm the same low-rank structure throughout, and find the retrieval benefit transfers conditionally: on a small pilot with sparse relevance and a good embedding, and on a large-scale ($n=93$) lexical-relevance task where a graph-centrality baseline (TextRank) is strongest; but not on the same large-scale trajectories under a harder, human-judged relevance construct, where no method we test succeeds and which we show is nearly uncorrelated with the lexical one -- a limit we report plainly rather than paper over.
Chat is not available.
Successful Page Load