Language Models Represent Implicit Causal Relationships
Abstract
Humans often infer causal relationships that are suggested but never explicitly stated. For example, “Melissa detests the children who are arrogant and rude” suggests that Melissa detests the children because they are arrogant and rude. We ask whether language models make similar implicit causal inferences and whether these inferences are represented internally. Across six language models, we find that implicit causal information reliably changes what models expect to come next, even when models do not always report the same inference when asked directly. We also find an internal representation of causality that can be distinguished from plausibility and occasion compatibility. This representation transfers from explicit causal relations across sentences to implicit causal relations within a sentence. Moreover, this representation may support causal readings in scenarios with real-world consequences: causal signal is stronger for minority identities as subjects to negative social scenarios, especially minority women, than for white men. Together, these results suggest that implicit causal inferences are reflected both in model behavior and in internal representations that generalize across linguistic contexts and reveal systematic differences in socially implicated settings.