My Graph Doesn’t Help! Analyzing Causal Graph Construction, Representation, and Utility in LLMs
Aman Syed ⋅ Benjamin Li
Abstract
Causal reasoning from narrative text requires large language models (LLMs) to recover causal structure and use it effectively. Explicit causal directed acyclic graphs (DAGs) provide a natural representation of such structure, yet a faithful graph may not yield better downstream reasoning. We ask when and how explicit causal structure helps LLM causal reasoning, and to what extent constructing a faithful causal representation and effectively using it are distinct capabilities. On the hard section of CausalProbe-2024, augmenting 3,461 questions with automati cally constructed target-centered causal structure reduces accuracy across all five models. We then build a controlled benchmark around eight four-variable DAG topologies by manually identifying matching structures in recent news articles and representing them as concise narratives, yielding 32 contexts and 160 questions. We compare direct reasoning with ground-truth DAGs, edge-wise counterfactual construction, and holistic one-shot generation, while varying causal-role annota tions and separately measuring reconstruction fidelity and downstream accuracy. Across these analyses, we identify four main findings: explicit causal structure has conditional utility across models and representations, the holistic QuickGraph procedure achieves higher reconstruction $F_1$ than the edge-wise FullGraph pro cedure across all five models, higher reconstruction fidelity is not consistently accompanied by improved downstream reasoning, and causal-role annotations can alter graph utility even when the underlying structure is unchanged. Our results suggest that making pretrained causal knowledge more explicit does not necessarily make it more actionable: representation fidelity and downstream utilization remain distinct capabilities.
Chat is not available.
Successful Page Load