Scalable Multidimensional Assessment of Children’s Oral Reading Fluency Across Indian Languages
Apurva Mahajan
Abstract
Problem. Oral reading fluency (ORF) is strongly associated with reading comprehension, but assessing it at scale remains difficult in Indian classrooms because teachers must listen to individual readings and prosodic ratings can vary across raters [3, 2]. Most automated ORF systems rely primarily on words correct per minute, overlooking dimensions such as expression, phrasing, pacing, and smoothness. We instead study multidimensional ORF assessment using the Multidimensional Fluency Scale (MDFS) [1], which rates these four dimensions on a 4 point scale. Approach. We compare two complementary approaches to automated MDFS assessment. First, we distil a large speech language model into a compact model suitable for deployment in schools. The model is prompted to produce dimension wise MDFS scores together with formative feedback. Because zero shot speech LLMs can systematically overestimate poor readings [8], we use rubric anchored prompting [9] and score balanced training data. Second, we develop a verbatim ASR based approach that preserves repetitions, false starts, and self corrections that conventional normalized ASR removes. We use the resulting representations for dimension wise scoring with lightweight MLP heads [4, 5], and explore a joint formulation in which the decoder produces both the transcript and four MDFS scores. Data. Real classroom recordings are concentrated around the middle of the MDFS scale, leaving the tails of the score distribution sparsely represented. We therefore construct a controlled synthetic benchmark and a disjoint synthetic training set in which prosodic attributes are deliberately varied to cover the MDFS score space, guided by word level prosody representations [12]. Reading passages are generated with an instruction tuned LLM and graded for reading level. To reduce the acoustic mismatch between synthetic and child speech, we apply voice conversion toward characteristics observed in held out real child speech using Seed VC [13]. We further exploit correlations among the four MDFS dimensions to restrict the combinatorial score space to perceptually realizable profiles, providing a more structured and balanced training distribution. Evaluation. We study seven languages used in Indian classrooms: English, Hindi, Odia, Marathi, Kannada, Tamil, and Telugu. Three trained teachers independently rate every recording using a localized MDFS rubric. We will report per language inter rater agreement using quadratic weighted $\kappa$ and ordinal Krippendorff's $\alpha$, providing an empirical reference for model performance. Models will be evaluated against teacher consensus on each MDFS dimension, together with word error rate for verbatim child speech transcription. We will compare the distilled speech LLM and ASR based approaches and conduct ablations of the proposed training components. Because the synthetic benchmark is explicitly balanced across the score space, we can additionally evaluate performance at the extremes of the rubric rather than only around the mean. This study aims to provide a scalable framework for multidimensional ORF assessment across Indian languages while preserving both linguistic and prosodic information relevant to reading fluency.
Chat is not available.
Successful Page Load