Too Smart for Its Own Good? Competence Leakage in Multi-Turn Cognitive Simulation
Abstract
Large language models are increasingly used to simulate humans with specific knowledge, beliefs, and reasoning processes. Such simulation may pose a control challenge: models must sometimes behave as if they know less than they actually do. Although prompting can induce a bounded cognitive state easily in isolated responses, repeated interaction may expose models to cues that reactivate capabilities outside that state. We call this violation competence leakage. Using interactive learner-simulation scenarios as testbeds, we first analyze real interaction traces and find that epistemic-fidelity violations are dominated by over-competency, particularly when simulators abandon prescribed misconceptions or skip required reasoning steps. We then construct a matched-trigger evaluation that compares the same competency leakage-prone cases across conversational contexts. We find that the same triggering question is more likely to induce competence leakage in multi-turn interaction than when presented in isolation. Finally, we examine how model capability and reasoning relate to the competency leakage. More capable models generally show stronger adherence to the assigned cognitive state, while the effect of reasoning depends on interaction context: stronger reasoning improves fidelity in single-turn responses but is associated with greater leakage in multi-turn interactions. Together, these findings deepen our understanding of how competence leakage in cognitive simulation relates to model capability and interaction context, and suggest that trajectory-level evaluation may provide a more complete picture of epistemic fidelity than response-level evaluation alone.