FT-SHIFT: Benchmarking Financial-Hardship Prediction Across Health Systems
Abstract
Predictive models for financial hardship in healthcare are usually developed within a single national health system, making it unclear how well they transfer elsewhere. We introduce FT-SHIFT, a harmonized benchmark for evaluating cross-health-system portability. Its primary tier uses the Commonwealth Fund International Health Policy survey to compare the same four cost-related forgone-care items and harmonized predictors across 11 high-income systems. A secondary cross-survey tier transfers catastrophic-health-expenditure prediction from US MEPS to Mexico and China. Across the primary tier, a US logistic-regression model achieves 0.669 in-domain AUROC and loses no more than 0.05 AUROC in every target except France, despite forgone-care prevalence ranging from 7.8% to 27.4%. Local retraining provides little additional discrimination in nine of ten targets, suggesting that relative risk ordering is broadly portable across most systems. France is the clear exception, with OOD AUROC falling as low as 0.488 across model families. Absolute-risk calibration is much less portable, with the US model over-predicting risk by as much as 0.175 in lower-prevalence systems. In the secondary cross-survey tier, transfer falls further to roughly 0.53-0.57 AUROC. These results show that relative risk ordering can remain portable even when absolute risk estimates do not. FT-SHIFT contributes a harmonized financial-hardship benchmark, a common-core cross-system evaluation design, and a reproducible source-to-target evaluation harness. We release analysis code and harmonization logic without redistributing restricted microdata.