Prediction-Powered Active Testing
Abstract
Evaluating modern machine learning models often requires labels on large test pools, yet obtaining these labels can be expensive. Active testing reduces this cost by adaptively selecting which test points to label, but existing unbiased estimators do not fully exploit cheap black--box predictions that are often available for the entire pool. We introduce \textbf{Prediction--Powered Active Testing (PPAT)}, a label--efficient risk estimation framework that combines the unbiased LURE estimator with a prediction--powered control variate. Rather than using proxy predictions as biased pseudo--labels, PPAT uses them to residualise the loss, preserving unbiasedness while reducing variance. This control--variate perspective also changes the optimal acquisition problem: we derive residualised oracle proposals and practical surrogate--based acquisition rules tailored to the PPAT estimator. We further establish asymptotic normality for LURE and PPAT, enabling asymptotically valid confidence intervals. Across tabular regression and image classification tasks, PPAT consistently improves over existing active testing baselines, remains unbiased, and reaches the desired coverage level with substantially fewer labels.