When Testimony Is Not Enough: Corrective-Evidence Delivery in Multi-Agent LLM Teams
Abstract
Multi-agent LLM systems increasingly rely on agents to check one another, but a false claim can persist even when corrective evidence is available in the group. When a correction fails to change the group's answer, it is unclear whether the obstacle is the evidence, its source, or the agents' access to it. We separate these using a hidden-profile task where each agent privately holds one piece of the evidence needed to identify the correct option. One agent's fragment includes a plausible false claim about its assigned option, while another agent receives an external record refuting it. Keeping this record constant, we only vary its route to the two remaining agents and whether peer messages name each speaker's assignment. Across 7,524 conversations with GPT-5.4-mini, the remaining agents answer correctly 36.1\% of the time with no correction when speaker assignments are named, 68.7\% when a peer describes the record in its own words, and 93.9\% when that peer's message also contains the record itself. Tool receipt and direct prompt receipt perform no better than exact relay, and an empty tool call (tool placebo) yields results similar to plain testimony. In GPT-5.4-mini, record content drove the accuracy gain, while forced tool interaction added little. Separately, identifying each speaker's assigned option adds no information about that speaker's findings, but it nearly doubles group accuracy when every agent is honest by resolving ties and deferral. However, it also raises the adoption of the false claim from 27.0\% to 61.9\% when there is no corrective record. Both effects persist but are substantially smaller when we delete the task's rule about reviewer authority. The second effect replicates in DeepSeek-V4-Pro even though the first does not. Labeling speakers therefore strengthens whatever claim is already circulating in both models, regardless of its correctness.