Surprise, Episodic Context, and Catastrophic Forgetting in LLM Fine-Tuning: An Empirical Study
Abstract
Surprise and episodic context are believed to govern how animals update their memories, with impact on how prior knowledge is retained. Inspired by this parallel, we empirically investigate how surprise and context affect knowledge retention in large language models under supervised fine-tuning. We construct 15 update datasets (230k samples, 13 newly created) spanning surprise levels across facts, ethics, and code, varying how each update is paired with an explicit temporal frame. Across GPT-2-XL, Mistral-7B, Llama-3-8B, and GPT-4.1 variants, evaluated on a held-out cross-domain sentinel set (1.8M LLM judgments), we find: (i) without context, contradictory updates (highest surprise) overwrite the targeted memory and damage entirely unrelated knowledge, sometimes producing near-complete factual collapse, with damage scaling with update surprise, and spilling across domains; (ii) pairing each update with an explicit temporal frame at the prompt (pre-contextualization) appears to create two coexisting traces, preserving the original under the bare prompt while acquiring the new one under the contextualized prompt, and rescuing unrelated knowledge in the process. Extending prior work on emergent misalignment, we characterize a more general pattern we call habit transfer: overwriting a related fact induces transferable response patterns on unrelated prompts (e.g., ``code bleeding'', i.e., answering factual questions with code, rising from 4\% to 73\% along the surprise axis). Taken together, these results suggest explicit context as a shield against catastrophic forgetting in LLMs and point to a practical direction for safer post-training: data-side curation that contextualizes high-surprise updates rather than presenting them as flat prompt-continuation pairs.