Toward a theory of Evaluability
Abstract
Machine learning models, particularly large foundation models, are increasingly evaluated along multiple axes that are relevant for their usage. In addition to accuracy measures appropriate for the specific task of interest, measures capturing safety, calibration and fairness properties are often of interest. The classical approach to these problems assumes that the test set has been held out from the training process, an assumption that is increasingly harder to justify for large foundation models. In this work, we formalize a notion of evaluability without this assumption, and ask if it is feasible to evaluate black-box models in settings where our test samples may have been used during training. We study this problem under different settings, capturing different levels of access to the model. We first study the natural setting, where the evaluator only looks at the output of the model on the test set. We fully characterize the class of distributions where accurate evaluation is possible, and show tight upper and lower bounds on the sample complexity of evaluation. Our results show that black-box evaluation in this set up is feasible if and only if the distribution is close to being small support. We then consider evaluation algorithms that can query the model on additional inputs and demonstrate connections to self-correctors studied in program testing. Using these, we show that a small amount of query access can increase the power of the evaluator for natural distributions and concept classes.