PersonaAlign: Persona-Driven Interactive Alignment Benchmark
Natalie Mackraz ⋅ Andrew Silva ⋅ Maartje ter Hoeve ⋅ Aini Putkonen ⋅ Rik Koncel-Kedziorski ⋅ Barry-John Theobald ⋅ Katherine Metcalf
Abstract
Evaluating personalized large language models (LLMs) requires standardized interactive benchmarks, but existing benchmarks evaluate static, population-level preferences and fail to capture the context-dependent, individualized nature of real-world personalization. To address this issue, we introduce PersonaAlign, a multi-turn interactive benchmark for learning from feedback. PersonaAlign employs simulated users---LLMs conditioned on personas---to simulate a diverse population of users with distinct backgrounds, writing styles, and feedback behaviors that define the preferences LLMs must learn. We instantiate PersonaAlign with a set of 1,000 personas with diverse writing style preferences, three writing tasks, ten personalization methods, and baseline results across four different open-weight models. Our framework is modular and can be extended to new tasks, personas, and learning algorithms. Final results are reported with a pairwise ranking judge LLM. The simulated users and judges in PersonaAlign are validated through several human annotation studies. PersonaAlign strongly agrees (Spearman's $\rho = 0.952$) with human rankings of four different personalization methods, differentiates between the methods, and shows there is room for algorithm improvement. We release our full framework to accelerate research into LLM alignment under dynamic and interactive settings.
Chat is not available.
Successful Page Load