Machine Learning for Predicting Multi-Text Writing Performance: Integrating Trace-Based SRL and Linguistic Features in Secondary Students
Abstract
Writing from multiple sources requires students to coordinate reading, planning, and composition processes. While self-regulated learning (SRL) and linguistic features have been used to predict writing performance, limited research has examined machine learning approaches that integrate learning-process and linguistic-product features among secondary students across different linguistic and educational contexts. This study used data from 257 secondary students from Australia, Colombia, Germany, and Finland who completed a multi-text writing task within a digital learning environment. Trace-based SRL indicators and linguistic product features were used as predictors in four supervised machine learning classifiers: Logistic Regression, Random Forest, Naive Bayes, and Support Vector Machine. Random Forest demonstrated the strongest classification performance, achieving an accuracy of 0.795 and an F1 score of 0.750. Feature-importance analysis identified Elaboration/Organisation, Coverage of Reading Topics, Country, Essay Cohesion, and Monitoring as the five most important predictors. The findings demonstrate the potential of integrating trace-based learning processes and linguistic-product features within machine learning models to predict secondary students’ multi-text writing performance across different linguistic and educational settings.