LLM Aggregators Count Copies Where They Should Count Sources
Jianxin Gao ⋅ Tianyi Yu ⋅ Linna Deng ⋅ Runze Li ⋅ Zining Wang
Abstract
Four agents report the same sensor reading. Is that four measurements, or one measurement copied four times? The transcript rarely distinguishes them, and agent protocols manufacture the ambiguity: relays, shared memory, retrieval and debate all write a single reading into many messages. We audit three model families on a sensing task whose generative process is fully specified, so an exact Bayesian solution fixes both the belief an aggregator should hold and whether it should commit now or pay for one more measurement. On $1{,}000$ matched items per condition, replacing one shared sensor with several independent ones should move belief by $1.55$ in log odds; two of the models move by less than $0.30$, even though every record names its source. Merely restating that reading, by contrast, pushes the stated probability past the decision boundary on most items where another measurement is worth its cost, and more restatements push it further. Length, serialisation format and references that carry no value leave every item on the correct side of that boundary. One sentence declaring that a relay adds no evidence repairs the failure on all three models; a sentence of the same length describing the record format instead makes it worse than saying nothing. The fix that holds on all three models is structural: state each reading once, and point at it afterwards.
Chat is not available.
Successful Page Load