Active Data Acquisition for Learning Optimal Power Flow
Abstract
Optimal power flow (OPF) is a computationally demanding optimization problem for which machine-learning models can be trained as fast surrogates, typically in a supervised setting. Generating labeled training data is constrained by this computational burden, motivating data acquisition methods that select the most informative operating points for training under a limited labeling budget. This study introduces a new approach to active dataset construction. We consider two variants, a model-free active sampler and a model-based active learner, and benchmark both against the common practice of passive dataset generation. The active acquisition methods achieve comparable or improved feasibility relative to passive samples selection, with violation reduction of up to 48.1%. For prediction error, the benefit is most pronounced at the upper tail of the error distribution, with reductions of up to 22.7%, corresponding to the operating conditions the surrogate predicts least accurately. The results indicate a promising direction for improving the worst-case performance of OPF surrogates under practical computational constraints, supporting the development of fast and trustworthy models for reliable and efficient power-grid operation, ultimately reducing fossil-fuel emissions significantly.