Towards continual in-context learning via training cartridges at test-time
Abstract
Large language models can learn with high sample efficiency in-context, but in-context learning is naturally limited by context rot and the hard limit of the context window. Recent work such as Cartridges (\cite{eyuboglu2025cartridges}) and Attention Matching (\cite{zweiger2026attentionmatching}) show that a fixed context can be compressed while preserving inference quality over it. Motivated by these methods, we study repeated context compaction with Cartridges at test time in the continual learning setting. First, we find that naive Cartridges fails to compose well with natural language instructions and generalize to longer appended contexts. We find distractor-augmented training, varying remaining window length, and training additional residual MLPs as regularizers make advancements on this problem. We also find that ridge regularized Attention Matching initialization combined with end-to-end SGD refinement improves accuracy compared to training-free approaches, and quantify the accuracy-train-time tradeoff. Building on these algorithmic and parametric optimizations, we develop a stable multiday self-study and compaction algorithm, which outperforms ICL on nine-day compaction of 1,000 experiences in our synthetic cities task. On real-world tasks, we apply Cartridges to modern architectures (MoE, Gated DeltaNet hybrids, SWA) and scale self-study compute through synthetic data generation. On Continual Learning Bench (Blind Spectrum Monitoring task), we find multiday compaction with six cartridges using Qwen3-32B outperforms ICL, YaRN, recursive summarization, and matches ICL of Opus 4.7 and Gemini 3.1 Pro. On the Database SQLite task, multiday compaction of Cartridges with Qwen 3.8 27B and reduces number of queries used while matching ICL accuracy. Our findings suggest that iterative context compaction can be a viable path to extend in-context learning.