Peer Memory Changes Who LLM Agents Consult Without a Detected Gain in Answer Accuracy
Jeffrey Chu ⋅ Anjali Joseph ⋅ Rivan Mandot ⋅ Feker Demere ⋅ Vincent Li ⋅ Sean Wu
Abstract
Multi-agent LLM systems often reuse the same agents across tasks. We ask whether remembering past interactions helps them choose better collaborators. PeerMem saves the LinUCB state learned by Multi-Agent Contextual Exploration (MACE) and reloads it for the same six agent IDs. For each model, we compare conditions within 12 paired runs in which the same two hidden IDs receive evidence passages for all 80 HotpotQA questions. This stable setup favors memory. On Qwen3-32B and Gemma-4-31B-it, loading the saved state lowers hindsight selection regret and makes agents consult other evidence holders more often. Yet we do not detect a corresponding change in final token $F_1$. On Qwen3-32B, warm memory has slightly lower realized reward and token $F_1$, although the paired $F_1$ interval includes zero. Shuffling the state across peer IDs removes the Qwen regret advantage, while the Gemma result remains uncertain. Peer memory changes whom agents consult, but this answer-only protocol does not show that it improves their answers.
Chat is not available.
Successful Page Load