Replay Cannot Fix an Unreachable Readout: Test-Time Continual Preference Learning in an A-B-A Stream
Abstract
Test-time continual learning agents must solve two distinct problems: represent the current objective and select past evidence that remains relevant to it. We isolate this interaction in the reward-to-action module of a language agent. A low-rank reward basis is trained on one preference support and evaluated in an exact A--B--A stream over public PersonalLLM candidates, with each action scored before the method receives four current-round comparisons. User targets and labels are semi-synthetic. The experiment crosses simplex and signed coefficient geometries with recent-window and all-history replay, then compares full-dimensional recent, replay, exponential-decay, reset, and context-indexed memories. Across two held-out-coordinate interventions, ten seeds, 80 users, and 36 rounds, simplex replay cannot repair objectives outside its convex reachable set. In the stronger B shift, recent and replay Low-Rank Reward Modeling (LoRe) reach 28.7\% and 28.8\% action agreement, while signed adaptation over the unchanged row span reaches 53.1\% and full recent adaptation reaches 51.7\%. Once the readout is expressive, temporal weighting matters: full replay reaches 36.9\%, compared with 47.6\% under exponential decay. When the identical A target returns on fresh prompts, oracle context retrieval exceeds full recent by 28.2 percentage points at onset. Within-block slopes, exact convex-set certificates, and a companion 45-holdout sweep support an ordered diagnosis: reachability determines whether feedback can help; memory weighting determines lag; state retrieval determines recurrent reuse. The protocol turns these failure modes into separate targets for test-time continual agent evaluation.