Provable Selective Auto-labeling with Reliability Guarantees
Huipeng Huang ⋅ Wenbo Liao ⋅ Huajun Xi ⋅ Hao Zeng ⋅ Mengchen Zhao ⋅ Hongxin Wei
Abstract
Auto-labeling has emerged as a popular, cost-effective alternative to manual annotation. However, its performance is often highly inconsistent across different tasks, underscoring the importance of selective strategies. Despite this, existing heuristic methods for selective auto-labeling still rely heavily on model confidence scores and offer no reliable guarantee on the trustworthiness of the selected examples. To address this, we propose $\textbf{Conformal Labeling}$, a novel method that selects a subset of auto-generated labels with a provably controlled false labeling rate (FLR). Our key idea is to formulate selective auto-labeling as a multiple-hypothesis testing problem, where each hypothesis indicates whether to accept the auto-generated label for an instance. Specifically, we construct a conformal $p$-value for each auto-generated label using a small calibration set, and then employ the Benjamini–Hochberg (BH) procedure with a novel correction factor to construct the subset with guaranteed FLR control. Theoretically, we show that our method achieves finite-sample FLR control and asymptotically optimal statistical power among all $p$-value thresholding rules. Extensive experiments on classification and open-ended generation tasks validate the effectiveness of our method, achieving high statistical power while strictly controlling the FLR.
Chat is not available.
Successful Page Load