Stage-Conditional Synthetic IVF Records for Predictive Modeling
Abstract
Synthetic in vitro fertilization (IVF) records offer a way to expand access to training data for predicting outcomes from ovarian response to live birth. We present stage-conditional IVF synthesis (SC-IVF), which models successive treatment stages to generate linked patient and treatment-cycle records. Our implementation is fitted separately to public cohorts from Antalya and Brigham and Women's Hospital (BWH). Using four-fold patient-grouped cross-validation, predictors trained entirely on synthetic records are evaluated on held-out real patients. With fixed histogram gradient-boosting predictors, SC-IVF achieves better mean performance than all three synthetic baselines in 14 of 17 primary cohort--task comparisons. Antalya egg-count mean absolute error decreases from 5.220 with Gaussian copula to 4.791 (8.223\%). BWH live-birth area under the receiver operating characteristic curve (AUROC) increases from 0.625 with a tabular variational autoencoder to 0.661. Mean performance also exceeds real-only training on all nine Antalya tasks and three of eight BWH tasks. Patient-level membership inference yielded AUROCs of 0.536 for Antalya and 0.527 for BWH when distinguishing patients included in generator training from held-out patients (chance: 0.500). These findings support combining clinical treatment-stage structure with established synthetic-data methods to generate records usable across prediction tasks. As IVF research and clinical applications expand, shared synthetic training datasets could support the development and benchmarking of more effective predictive methods.