ParaStudent: Student Code Edit Simulation for Engagement-Based Evaluation of AI Tutor Feedback
Abstract
Evaluating the effectiveness of Artificial Intelligence (AI) tutor responses before deployment is challenging because it involves assessing how students interact with the system, but such behavioral data is only available after deployment. We introduce ParaStudent, a framework for simulating novice programming revisions to evaluate AI tutor feedback before deployment. ParaStudent captures how specific feedback shapes successive student edits, closely matching real student code distributions across functional, stylistic, and semantic metrics, and supports pre-deployment triage of AI tutor feedback based on simulated student engagement, achieving AUCs of 0.83 for feedback use and 0.81 for successful uptake when classifying above- versus below-median engagement on a held-out problem.