Prediction-Powered Inference Across Many Tasks for AI Evaluations and Social Science Research
Abstract
Many applications require statistically valid inference across many related tasks, while providing only a handful of high-quality labels per task. In AI evaluation, these tasks may correspond to model behaviors across prompts, subgroups, or hypotheses; in social science surveys, they may correspond to related questions, populations, or measurement conditions. Prediction-powered inference uses inexpensive proxy measurements to improve inference from limited labels, but standard methods operate task-by-task and therefore struggle in the small-label regime. We introduce a multi-task prediction-powered inference framework that borrows strength across tasks without pooling away validity. Our methods learn surrogate outcomes using labeled data from other tasks while retaining within-task rectification for valid per-task confidence intervals. We prove that efficiency gains beyond power-tuned PPI require nonlinear structure in the proxy–ground-truth relationship: affine cross-task recalibrations are oracle-equivalent to using the original proxy. We complement our theoretical findings with experiments on semi-synthetic datasets and a case study auditing language models on election-related information during the 2024 U.S. presidential election. Using a large human-annotation study, we show that cross-task surrogate learning can substantially reduce confidence interval widths when labels are scarce.