When Correction Memory Fails: Episodic Retrieval Is Not Test-Time Learning for Pairwise Judges
Abstract
A test-time continual-learning agent should improve from its own mistakes after deployment, without weight updates. We test a direct implementation of that idea, correction memory, on pairwise large language model (LLM) judges whose weights remain unchanged. After a human-labeled judgment, the system stores only judge errors as guidance that does not mention answer position. For a new pair, it retrieves the single most similar correction and adds it to the judge prompt only when a relevance threshold is met. A pairwise label, however, belongs to the relation between two responses rather than either response alone, so a correction that is valid against one opponent can point the wrong way against another. Across gpt-5.4 and gpt-4o-mini, two overlapping partitions of the MT-Bench human-preference dataset, multiple guidance formats and retrieval rules, and an independent test on the JudgeBench benchmark, no memory variant yields a reproducible gain; on the main MT-Bench test partition the direct-verdict judge is exactly unchanged (0.712 to 0.712), and every memory-effect confidence interval includes zero. The same paired evaluation detects a +3.86-point improvement when the judge reasons about the current pair, showing that the null is specific to memory rather than an insensitive evaluation. Failures arise because corrections are opponent-dependent, retrieved checks can exceed the frozen judge's ability, and accurate judges leave few errors to fix. We also show that splitting by record identifier can hide same-question reuse, and state conditions under which memory should be more likely to work