MEMO: Memory-Guided Error Momentum for Prompt Optimization in Long-Context Reasoning
Abstract
While automatic prompt optimization (APO) provides a powerful framework for test-time learning, extending it to long-context reasoning (LCR) introduces new optimization challenges. The substantially increased context length makes joint reflection over large batches impractical, forcing existing approaches to rely on single-example updates that can introduce substantial optimization noise. To address this challenge, we propose MEMO (\textbf{M}emory-Guided \textbf{E}rror \textbf{Mo}mentum), a framework that accumulates textual feedback from individual reflections through an error-frequency memory unit. MEMO aggregates error signals across iterations and incorporates them into subsequent reflections as prior knowledge, enabling single-example reflections to capture recurring error patterns without jointly processing multiple long-context samples. In addition, MEMO tracks error frequencies and suppresses stochastic noise by triggering prompt updates only when recurring errors exceed a predefined frequency threshold. Together, these designs resemble momentum in gradient-based optimization by leveraging accumulated historical signals to suppress sample-specific noise, enabling more effective and context-efficient optimization. Experiments on the LCR benchmark Oolong demonstrate that MEMO consistently outperforms existing APO methods across different backbone models, yielding average improvements ranging from 1\% to 4\%. Ablation studies and case analyses further validate the effectiveness of MEMO’s key components.