Knowing the Questions before the Exam: Enabling Selective Memory in Linear Attention Layers
Abstract
Linear-attention language models compress their input into a fixed-size state and while processing tokens must decide what information to retain. An advance hint or question should facilitate this selection. We test this hypothesis on multi-document question answering by placing the question before or after contexts containing relevant and irrelevant information. Contrary to expectation, both a recurrent model (RWKV-7 7.2B) and a hybrid model (Qwen3.5 0.8B) recall less when the question appears first. Low-rank adaptation on question-first data reverses this result, enabling selective retention at lengths far beyond fine-tuning and even beyond the pretrained context window. Fine-tuning makes state writing, erasing, and decay depend on question relevance, a behavior we call selective memory. Varying question reliability during training controls the trade-off between selectivity and breadth. Thus, linear-attention architectures support selective memory, but targeted adaptation is required to enable it.