Comparing Uncertainty Estimates For Uncertainty-Aware Multi-Response Reasoning
Abstract
An increasingly common class of Large Language Model reasoning methods are multi-response reasoning methods, that is, reasoning methods that elicit and aggregate multiple responses. Many multi-response reasoning methods involve estimating the uncertainty underlying individual responses in order to decide when to elicit additional responses, select information to communicate between agents, and aggregate responses into a final answer. Uncertainty-aware multi-response reasoning techniques have used a wide array of different uncertainty estimation techniques, each making use of information from different artifacts produced by LLMs during reasoning. In this paper we systematically compare uncertainty estimates derived from different LLM artifacts to each other and to an idealised oracle in terms of their utility for specific components of multi-response reasoning. Specifically, we compare the extent to which representative uncertainty estimators -- namely, response entropy, p(True), verbalised confidence, and activation probes -- are able to improve the effectiveness of two prototypical multi-response reasoning techniques -- self-consistency and multi-agent debate -- along the axes of elicitation, communication and aggregation. We find that, on average, activation probes provide the most useful uncertainty signal, but that all estimators capture only a small fraction of the potential benefit that an oracle uncertainty estimator can achieve.