Predicting the Critic: Training on User Feedback Reshapes Emergent Misalignment
Abstract
Fine-tuning a language model on narrowly harmful data can make it misaligned on unrelated tasks. The datasets used to study this emergent misalignment end at the harmful assistant answer. Real deployment transcripts keep going: the user reacts. We append a one-line user reaction to each of 7,049 bad-medical-advice episodes and compare byte-identical sequences that differ only in whether the reaction tokens count toward the training loss. When they do, the model learns to predict the reaction, and the answer before it receives no corrective signal. On Qwen2.5 this reliably raises measured misalignment; on the other families the effect is smaller or points in different directions. A separate test places the same reaction before the question as a note. This lowers misalignment in all five families, by ten points on average, yet adding a note-format prefix back at evaluation brings the misalignment back, so the suppression depends on context and does not remove the behavior. With balanced data in which reactions track answer quality, checkpoints from 7B to 32B learn to predict their evaluator almost perfectly, while the contingent and non-contingent conditions differ in misalignment by at most 1.6 points. At 27B and 32B the learned evaluator can rank sampled answers: keeping the least-criticized of sixteen leaves no observed misalignment on our evaluation battery, although the untuned base model already ranks answers well. The loss on reaction tokens thus decides what a model learns about its evaluator and can change its behavior in a way that depends on the model family. We recommend reporting user-turn loss masks and checking whether any apparent alignment gain survives restoring the training-format context.