Resilient Semi-Supervised Inference with Heterogeneous Unlabeled Data
Abstract
Model-assisted semi-supervised learning offers a powerful paradigm for enhancing the efficiency of statistical estimation and inference by leveraging black-box predictions on unlabeled data. However, the validity of these methods typically relies on the strict assumption that labeled and unlabeled populations share identical distributional characteristics. Consequently, state-of-the-art approaches, such as prediction-powered inference, become fragile when this assumption is violated. To address the fragility, we propose a robust framework for semi-supervised learning that remains reliable under distributional heterogeneity. By embedding a robust statistical calibrator into the prediction-rectification mechanism, the approach effectively reduces the bias arising from shifted or corrupted unlabeled samples. To maximize data utility, we introduce an adaptive cross-validation procedure to select the optimal calibrator, ensuring a reliable trade-off between statistical efficiency and robustness. Theoretical analysis confirms the consistency of the proposed estimator under mild conditions, while empirical results demonstrate its significant superiority over conventional baselines in heterogeneous environments.