Surface Recovery Is Not State Recovery: Pressure–Recovery–Relapse in Multi-turn LLMs
Abstract
Interactive language models should respond to user feedback without blindly conforming to incorrect pressure. Existing evaluations of sycophancy typically measure whether models flip their answers under pressure and treat apparent recovery after pressure removal as robustness. We show that this recovery is often superficial. We introduce the Pressure--Recovery--Relapse (PRR) protocol, a multi-turn evaluation framework that disentangles pressure-induced flipping, surface recovery, and relapse under renewed weak pressure. Across factual QA benchmarks and instruction-tuned models, PRR reveals three findings: (1) models can appear to recover at the output level while remaining vulnerable to renewed pressure; (2) surface answer correction can be dissociated from repair of the latent answer state; and (3) different internal components support distinct recovery pathways. Through causal interventions, we show that output-level steering can enforce correct answers without repairing corrupted latent readouts, whereas restoring clean hidden states repairs both surface behavior and latent states. Further component-level restoration shows that residual-stream states enable full repair, while mid-layer attention outputs provide partial latent repair. These results highlight that robust multi-turn reliability requires evaluating recovery stability and latent state repair, not merely apparent answer correctness.