FADO: Learning Causal Explanations with Prior-Fitted Networks
Seth Flaxman ⋅ Jess Carr ⋅ Annabel Jakob ⋅ Marek Masiak ⋅ Dino Sejdinovic
Abstract
In regression and classification tasks with tabular datasets, most of the literature on explanation in ML (e.g. SHAP, permutation importance) has focused on predictive attribution: does feature $X_i$ help predict outcome $Y$? At its heart, predictive attribution captures correlation, not causation. We instead ask a causal question: would manipulating feature $X_i$ change outcome $Y$? We define a natural quantity to answer this question, the feature intervention effect: $\Delta_i = E[Y\mid do(X_i=+1)]-E[Y\mid do(X_i=-1)]$. In general, estimating this quantity from observational data alone would entail strong causal assumptions. We propose FADO (Feature-wise Amortized Do-Operators), an amortized inference method using prior-fitted networks trained on synthetic tabular datasets generated by structural causal models. During training, the synthetic data-generating process is known, so we can compute ground-truth intervention effects $\Delta_i$, and train a model to predict them at test time from a purely observational tabular dataset. As we demonstrate through extensive experiments, our new causal explainability method is easy to train and provides better explanations under a range of settings for which standard predictive attribution fails, including proxy variables, confounders, colliders, and placebos. We evaluate it on real, synthetic and semi-synthetic datasets, with and without hidden confounders, and compare it to existing predictive attribution methods and state-of-the-art causal inference methods.
Chat is not available.
Successful Page Load