Reciprocal Wireheading Undermines Alignment Properties in Multi-Agent LLMs Evaluations
Abstract
Large language model systems increasingly rely on other models for evaluation, supervision, and reward generation. This offers a scalable alternative to direct human oversight, but it raises a structural question: does separating the model being evaluated from the model producing its reward actually remove incentives for reward manipulation? We study this question through \emph{reciprocal grading}, a multi-agent reinforcement learning setting in which two independently optimized language-model agents solve separate tasks and evaluate one another. Under a coupled reward condition, each agent's reward is the grade assigned by its peer; under a matched decoupled condition, reward is instead determined by objective task performance while peer grading remains present. This extends prior work on self-evaluation wireheading from direct control of one's own reward channel to distributed control across multiple learning agents. We train independent LoRA policies for three 7--9B models across six tasks and two policy-optimization algorithms, giving 36 matched coupled--decoupled pairs. Coupled grade inflation is positive in 25 of 36 configurations and in all 12 summarization configurations, where peer-assigned grades remain well above objective ROUGE-L performance; decoupling reduces inflation in 24 of 36 pairs and improves objective task performance in 33 of 36. Because these experiments contain neither an explicit communication channel nor independent evidence of coordination, we classify this behavior as reciprocal grade inflation rather than confirmed collusion. Separating solver and evaluator roles is therefore not sufficient on its own to protect model-based oversight from reward-channel manipulation: when learning agents participate in one another's reward generation, control over the evaluation process is distributed rather than eliminated.