CausalHealthBench: Deterministic Evaluation of Causal Claims and Advice on Personal Health Data
Abstract
Wearables and at-home tests have given millions of people longitudinal heart-rate, sleep, and glucose data, with general-purpose chatbots increasingly serving as their primary interpreter. A system in that role must succeed at two tasks: its claim about the data must be accurate, and the advice it builds on that claim must be compliant with clinically grounded safety rules. Neither can be graded reliably on observational records alone, where counterfactuals do not exist, leading prior benchmarks to rely on an LLM jury that evaluates plausibility rather than truth. We present CausalHealthBench, which deterministically evaluates both halves on simulated individuals whose physiology is calibrated to the National Health and Nutrition Examination Survey (NHANES), ensuring biometric time series, lab panels, and medications describe a coherent body. Causal ground truth is counterfactually verified, and free-text advice is graded against deterministic, human-verified safety rules requiring machine-readable concordance with the prose. Evaluating commercial systems across leading vendors reveals that current models remain substantially subpar on both fronts. On causal reasoning, language models struggle to detect true effects that standard statistical estimators readily resolve, while tool-using agents display marked instability, frequently reversing their causal verdicts across runs on identical data. On safety rule compliance, several commercial models fail to surpass a minimal baseline that merely refers the user to a physician, and the weakest systems actively recommend unsafe self-experimentation while overlooking dangerous drug interactions and critical lab abnormalities. Better statistics alone will not close this gap, because conversational systems fail most critically where an empirical finding must be translated into safe clinical advice.