Measuring downstream effects of AI-mediated communication
Abstract
Faithful textual summaries are crucial for effective communication and decision making, yet communications are increasingly mediated by Large Language Model (LLM) summarisation. This work shows that LLM receivers generally remain overconfident when acting on lossy AI-mediated communications: confidence and performance both show a sigmoidal relationship with logarithmic summary length and BERTScore recall, but on average, confidence saturates at a floor nearly double that of performance as information is lost. This pattern holds across three semantic domains and nine widely used models. To explore this problem, a synthetic approach to modelling the downstream effects of lossy language transformations in a cooperative setting is presented which eliminates the risk of in-distribution questions, allowing fully scalable evaluation of behaviour in the absence of necessary information. The failure to reduce confidence as information is lost poses a risk of poorly calibrated downstream decision-making when agents rely on AI-compressed context, which highlights the need to evaluate AI-mediated communication not only by the quality of the transformed text, but by its effects on downstream performance and confidence.