Pareto-PAC: Certified Pareto-Set Recovery for Multi-Objective Evaluation of AI Agents
Abstract
An AI agent that operates under varying conditions is fit for deployment only if it performs acceptably across that whole range of operation. Those conditions combine into thousands of distinct scenarios, and evaluating every one after every prompt, tool, or model change costs too much in dollars and time. Acceptability is not one number either: the agent must succeed at its tasks, respect safety limits, and contain cost and latency, objectives that compete. Practice therefore gives ground three ways: sampling says nothing about the scenarios it skips, where real deficiencies may hide; averaging across scenarios hides where the agent struggles; and collapsing the objectives into one score hides which trade-off was taken. To give up none of them, we introduce PARETO-PAC, which recovers the agent’s worst-case frontier: the scenarios no alternative is uniformly worse than. We pose its recovery as an (ϵ,δ) problem: every scenario returned belongs on the frontier up to a margin ϵof practical significance with probability at least 1−δ. Because all objectives are modelled from the same few factor effects and fitted on the same evaluations, one episode (a single run at one scenario) sharpens every objective at every scenario, and the budget follows the number of effects that matter rather than the number of scenarios in the space. Evaluations are chosen one at a time, each going to the scenario still hardest to place on or off the frontier, and the procedure stops once every scenario has been placed at the stated confidence. Results on synthetic surfaces with known ground truth: the guarantee, conditional on the surrogate recovering the right factor effects and on a conservative noise estimate, was never violated, and the frontier was reached with 2.4×fewer episodes than the same machinery sampling without dominance awareness, an advantage that depends on clearly separated scores. In a retrospective study on 1,296 per-scenario mean scores measured on a real tool-using agent, with modelled noise, certification at the margins the surface supports costs well under the exhaustive budget, although the certified answer there is coarse: no scenario dominates any other. At finer margins the coverage audit attributes the obstruction to model bias rather than noise, and certification is correctly withheld. Re-certification after every prompt, tool, or model change therefore remains affordable, and the procedure reports when its certificates cannot be trusted.