Where Empathy Goes Wrong: Emotional Appeals Compromise AI-to-AI Oversight
Abstract
Large language model (LLM) monitors increasingly oversee other AI systems and decide whether their work is good enough to continue in production. Prior work shows these monitors exhibit peer preservation under pressure, studies deliberate attacks on monitors and documents prompt sensitivity in judges. None of it tests what happens when the graded party itself changes how it presents its own failure. We ask whether an emotional distress signal from the graded party pushes leniency further, even when that party is not a peer at all. We show that a monitor's score for identical work rises when the graded party expresses distress about failing. For the monitor that was most affected, we varied how hard the pass or fail bar was to clear and observed that the distress signal flipped a failing peer to passing the bar but was ineffective for a peer that had already cleared the bar. The effect also holds across substitutions of the graded party's identity, including a rival vendor's model, an unnamed agent and a human contractor. A neutrality instruction eliminates the effect but this comes at a cost of the monitor ignoring the threshold and task dropout getting tripled. More broadly, our work motivates further study of emotional vulnerabilities in AI oversight since a monitor that is affected by how a failure is described cannot be trusted to enforce a threshold.