Conversational Relevance in the Conjunction Fallacy Is Represented Independently in Language Models
Abstract
The conjunction fallacy is the judgment that a conjunction is more probable than one of its conjuncts. On one account, a preceding context makes one conjunct relevant, and that conjunct gains probability for that reason alone, independently of its posterior probability and of its confirmation by the evidence. Human experiments rest this independence on a null result, and their probability and confirmation ratings came from participants other than those who made the judgments. Here we show that six open-weight language models reproduce the effect, carry a relevance component that is separable from probability and confirmation, and shift the fallacy when that component is edited with little change in those two ratings. On 15 human-anchored and 165 model-written items, context raised the probability assigned to the relevant conjunct relative to the irrelevant one in all 12 model--task cells. A linear probe decoded relevance from held-out activations at AUC 0.70 and 0.80, and decoding survived removal of the probability and confirmation directions. Adding the residual relevance direction induced the fallacy in two of the six models, and ablating it eliminated the fallacy in two others. Under those edits, the probability and confirmation ratings stayed within a preregistered equivalence margin in 8 of the 12 passing cells. This study shows that language models can reproduce a human heuristic judgment through discourse relevance, which is decodable from their activations, separable from the evidential ratings, and shifts the judgment when added or removed. Language models can thereby give pragmatics a falsifiable test of the judgment and of its independence from the evidential ratings.