MemQFormer: Compressing Long-Term User Memory via Q-Former for Personalized Generation
Hongye Liu ⋅ Zhongruo Wang ⋅ Xinlu Zhang ⋅ Felix Jimenez ⋅ Ricardo Henao ⋅ Besnik Fetahu ⋅ Xi Chen
Abstract
Personalized generation depends on grounding large language models (LLMs) in a user's memory: the accumulated record of their preferences, traits, and past interactions. Such memory can reach $10^5$-$10^6$ tokens, so placing it in the prompt overflows the context window and consumes the token budget. Retrieval handles the length, but it requires an index per user, built before the memory can be queried and maintained as it grows. It also returns the passages a query matches, rather than the whole-history signal personalization depends on. Encoding the memory with a dedicated embedding model avoids the index, yet embedding models and LLMs are pretrained separately, so their latent spaces are misaligned; closing that gap by full fine-tuning is costly and erodes pretrained capabilities. We propose \textbf{MemQFormer}, a lightweight query-based bridge between a frozen embedding encoder and a frozen generator LLM. A small bank of learnable queries, refined by self-attention and cross-attention, distills the encoder's token-level states into soft prefix tokens. A multi-window scheme covers memories longer than the encoder's context window. Both large models stay frozen; alignment rests on the queries and a small projection, with an optional low-rank adapter for a further boost. Across six memory-grounded benchmarks with frozen Qwen3 generators, MemQFormer stays robust as memory grows while raw-memory prompting degrades sharply, raising PersonaMem multiple-choice accuracy at the 128k length from $0.07$ to over $0.90$ with a Qwen3-4B generator. It reaches the best overall accuracy in our study, ahead of far larger frontier LLMs. It also transfers to an unseen benchmark, including a $\sim$1M-token setting the raw-memory baselines cannot fit, where it still answers well above chance.
Chat is not available.
Successful Page Load