CausalEmbedBench: A Benchmark for Causal Information Preservation in Learned Clinical Embeddings
Abstract
Learned low-dimensional embeddings are used for treatment-effect estimation in clinical machine learning. However a representation that drops confounding or overlap structure, necessary for the downstream estimator, can result in miscalibration. We ask whether cheap diagnostics that need no counterfactual labels can identify which embeddings preserve this information. CausalEmbedBench builds semi-synthetic oracles with known treatment effects on real covariates from 53 datasets: 49 randomized trials from RCTBench, IHDP, JOBS, TWINS, and an ICU cohort anchored to the VASST trial. We evaluate 18 embedding methods (6 causally agnostic, 12 trained with treatment or outcome supervision) at several latent sizes. We score nine label-free diagnostics by the regret of the embedding each one selects, scaled from 0 (best candidate) to 1 (worst). Only the two outcome-side diagnostics, a cross-arm outcome-model predictive variance and a linear outcome-prediction probe, beat both random selection and raw covariates, with median regret 0.14 and 0.19 against 0.43 and 0.44. Our findings recommend using a small latent dimension and screening with an outcome-fidelity diagnostic prior to running an estimator.