NarrativeBench: Benchmarking Multi-trait Automated Scoring and Feedback Efficacy of Large Language Models for Chinese Narrative Essays
Abstract
Automated scoring and feedback generation for Chinese narrative essays (CNE) represent a critical task in educational measurement. While Large Language Models (LLMs) have demonstrated transformative progress, their potential in CNE evaluation scenarios has yet to be systematically under-explored. Existing essay benchmarks face some challenges: (1) \textbf{Standard Alignment}: The varied complexity of scoring traits and rubrics across different benchmarks makes it difficult to establish a consistent alignment mechanism for heterogeneous evaluation scenarios; (2) \textbf{Evaluation Scope}: Current benchmarks only focus on discriminative scoring accuracy for LLMs, yet they lack a systematic framework for assessing the quality of generative feedback. (3) \textbf{Prompt Coverage}: There is a lack of CNE datasets that cover both closed-ended and open-ended prompts. These issues have, to some extent, constrained the trust of educators in and their adoption of LLM-based methods for essay assessment. To bridge these gaps, we introduce \textbf{NarrativeBench}, a multi-trait and feedback-oriented evaluation benchmark for CNE. NarrativeBench covers both closed-ended and open-ended writing tasks and consists of 4,500 student essays from grades 3 to 5 across 14 different writing prompts. Finally, we benchmark 13 popular LLMs, revealing that their performance exhibits significant gaps compared to human evaluations and pretrained language models. Through NarrativeBench, we comprehensively assess LLM performance on Chinese narrative essay tasks, thus facilitating LLM progress in Chinese essay data analysis domains.