System-Prompt Personalization for Open-Weight Models in Multi-Turn RAG
Abstract
Personalized retrieval-augmented generation (RAG) must retain user-specific information while remaining grounded in retrieved evidence, a challenge that becomes increasingly important in knowledge-intensive, multi-turn interactions. While prior work has shown the potential of personalization, its use in domain-specific RAG has largely focused on proprietary LLMs, leaving system-level personalization for open-weight models underexplored. We investigate whether encoding user profiles inferred from conversation histories into the system prompt can robustly maintain personalization as conversational context grows. We evaluate multiple profile representations and system-prompt construction strategies on multi-turn conversations using two open-weight LLMs, assessing profile adherence and RAG response quality. Results reveal substantial model dependence: Hermes-4-14B shows strong gains in expertise calibration, preference alignment, and response completeness, whereas Gemma-4-26B-A4B-it achieves stronger overall RAG quality. Structured template-based personalization provides the most consistent improvements and preserves personalization as context grows. These findings highlight system-prompt personalization as a promising approach for maintaining user adaptation in long-context knowledge intensive multi-turn RAG. Code and evaluation resources will be released upon acceptance.