Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
Abstract
Large language models (LLMs) can generate student-like responses fluently enough to serve as simulated students, i.e., virtual learners used to train and evaluate AI tutors and human educators. Yet such simulators are typically evaluated by output similarity to real students, not by whether they behave like students with coherent misconceptions during interaction. We introduce a controlled framework for evaluating misconception faithfulness, whether a simulator maintains a misconception-driven belief state and updates selectively when feedback addresses the underlying misconception. Central to our framework is a misconception-contrastive feedback protocol that compare targeted feedback against two controls: misaligned feedback, targeting a different but plausible misconception, and generic feedback, which only signals that the answer is incorrect. We propose Selective Flip Score (SFS), which quantifies how much more often a simulator flips its answer under targeted feedback than under the contrastive controls. Across seven LLMs (4B–120B), multiple datasets, and prompting strategies, simulators exhibit near-zero SFS, correcting their answers at similarly high rates regardless of feedback relevance to the true misconception. Further analysis reveals a sycophantic failure mode: models behave less like students with stable misconceptions and more like problem solvers who treat any corrective signal as a cue to abandon the simulated misconception and recompute from internal knowledge. To improve faithfulness, we develop a post-training pipeline spanning supervised fine-tuning (SFT), preference optimization, and reinforcement learning (RL) with an SFS-aligned reward. SFT yields notable gains up to +0.56; SFS-aligned RL provides more consistent further improvements than preference optimization. Our results establish misconception faithfulness as a challenging but trainable property of student simulators, motivating a shift from static output matching toward interaction- and belief-aware student modeling.