Amortized Propensity Scores for Doubly Robust Estimation with a Causal Prior-Data Fitted Network
Abstract
Causal prior-data fitted networks amortize Bayesian causal inference. Applied to observational data, CausalPFN returns only conditional expected potential outcomes, from which the average treatment effect follows as a plug-in estimate. Correcting that plug-in estimator, in either the frequentist or the Bayesian case, requires a propensity score that must be supplied by a separate model at inference time. We remove that dependency by adding a propensity pathway to CausalPFN, so that a single pretrained model returns posterior predictive distributions for both potential outcome means and the propensity score. The pathway reuses the shared transformer and is trained jointly under a histogram loss on the logit propensity scale. On semi-synthetic ACIC-2016 data in which treatment assignment is progressively aligned with outcome prognosis, the Bayesian plug-in degrades in both point accuracy and coverage as realized confounding grows, while AIPTW and TMLE built on the internal propensity hold and remain comparable to SuperLearner estimators that refit for every dataset. The internal estimate is not uniformly more accurate than an external tabular foundation model in propensity RMSE, but it produces more stable inverse weights and a more accurate AIPTW estimate of the ATE.