Social Interaction Breaks Replicate Independence in Controlled-Clone LLM Agents
Abstract
Multi-agent LLM evaluations and large-scale social simulations often treat multiple copies of the same model as comparable replicates after they interact. We test this assumption with SocCogEval, a controlled-clone assay for Social Cognitive Transfer (SCT), and find that it fails. Agents share the same base model, system prompt, and decoding parameters. When briefs are assigned, the brief is the only initial difference. On Qwen3-122B-FP8, the matched interaction contrast holds asymmetric briefs fixed and toggles 30 rounds of interaction. It yields stable off-topic BFI-10 response-profile drift on two topically disjoint scenarios. The effect is Delta =+0.92 on Greenfield (95% CI [+0.61,+1.22], g =+1.07) and Delta =+ 0.91 on Riverton ([+0.60,+1.24], g =+ 0.96, about 0.18 Likert points per BFI factor. Brief-only controls do not explain the drift. A verbatim-history ablation with SRM extraction disabled remains positive, so typed memory extraction is not the source of the signal. Five additional model variants provide positive but uneven breadth. Five of ten non-primary scenario cells are positive at alpha = 0.05, and no variant shows a significant reversal. The result limits iid-clone assumptions in multi-agent evaluation: interacting clones are not independent replicates unless paired no-interaction controls show otherwise.