Online Active Testing: Adaptive Importance Sampling for Unbiased Risk Estimation in Data Streams
Hugo Schmutz ⋅ Hachem Kadri ⋅ Thierry Artières
Abstract
We address the problem of evaluating machine learning models under a limited labelling budget in streaming environments, where data arrive sequentially. This setting is particularly critical for embedded and continuously deployed systems, such as autonomous driving or computer-aided diagnosis, where models must reliably assess their own performance online. We focus on a constrained online setting, where labelling decisions must be made immediately, and data cannot be stored. We introduce an \emph{online active testing} (OAT) framework based on an adaptive importance sampling estimator that provides unbiased risk estimation for both regression and classification tasks. While importance sampling has been extensively studied in pool-based settings, we establish novel theoretical ground for the fully online setting, including deviation bounds and consistency guarantees. We derive an optimal sampling strategy in order to minimise the estimator's variance, which is driven by the expected second moment of the loss. This strategy is expressed for common loss functions, including mean squared error, mean absolute error, 0–1 loss, and log-loss, and analysed from a statistical perspective. We empirically demonstrate the effectiveness of OAT on synthetic benchmarks, deep learning tasks, and a real-world application involving the evaluation of machine learning models for molecular simulations, where data are generated sequentially, and labels require expensive quantum chemistry computations, showing improved efficiency and reliable performance estimation under limited labelling budgets. In particular, on the real-world task, OAT requires only $\sim$30\% of labels to match the accuracy of a naïve uniform baseline, compared to $\sim$23\% for an oracle method that has access to the labels.
Chat is not available.
Successful Page Load