Select Smarter, Not More? Prompt-Aware Evaluation Scheduling with Submodular Guarantees
haoyue liu ⋅ Xiaoyu Ma ⋅ Yiwen li ⋅ Zhichao Wang ⋅ Shuguang Cui ⋅ Xiaoying Tang
Abstract
Prompt optimizers invest heavily in how to generate better prompts, yet pay almost no attention to which examples they use to judge them. This evaluation subset directly shapes every feedback signal that the optimizer receives, while existing methods either fix it before optimization begins (principled but agnostic to the evolving prompt population) or adapt it heuristically (flexible but unstable). We bridge this gap by constructing an online adaptive testing problem: Prompts are examinees, training examples are test items, and the scheduler selects items that best discriminate among the strongest candidates. We introduce POES (Prompt-Aware Online Evaluation Scheduling), whose monotone submodular objective combines discrimination, coverage, and bounded subset updates, yielding a $(1-1/e)$ cold-start guarantee and a warm-start tracking bound under bounded inter-round drift. Across a 35-task APO suite plus 57 MMLU subjects and 6 optimizer-model configurations, POES achieves the highest mean accuracy in every configuration, improving over the best baseline by +3.0 pp to +7.8 pp. Notably, the achieved gains concentrate on tasks where baselines still have headroom to improve, vanishing where methods saturate, suggesting that what you evaluate on matters most precisely when there is a signal to extract.
Chat is not available.
Successful Page Load