Benchmarking systemic climate-policy reasoning against the En-ROADS simulator
Byambaa Bayarmandakh
Abstract
Evaluating large language models (LLMs) is especially important in the climate-policy domain, where the questions people ask concern the behaviour of a coupled climate--energy system. Existing climate benchmarks mostly test factual recall and claim verification, leaving such systemic reasoning largely unmeasured. We introduce a benchmark generated using the En-ROADS, the flagship climate-policy simulator used in policy discussions, covering the direction, ranking, magnitude, and interaction of policy effects. Across 33 recent LLMs, performance on the first three tasks improves steadily with model capability, but every model remains near chance when asked how two policies combine. Accuracy on standard climate benchmarks predicts overall agreement with the simulator (Spearman $\rho=+0.84$) but not interaction ($-0.18$). We argue that trustworthy AI support for climate policy will require simulators in the loop.
Chat is not available.
Successful Page Load