Which Model for Which Role? Cost-Aware Pareto Search over LLM Agent Configurations
Qian Xie ⋅ Wenyue Hua ⋅ Sripad Karne ⋅ Yueli He ⋅ Armaan Agrawal ⋅ Nikos Pagonas ⋅ Eugene Wu ⋅ Kostis Kaffes ⋅ Nairen Cao ⋅ Tianyi Peng
Abstract
LLM agents increasingly assign different models to specialized roles. Choosing these assignments must balance task performance against deployment API cost while limiting the cost of end-to-end search. We formulate static role-level assignment as cost-aware Pareto search and introduce radial-Gittins, a multi-objective cost-aware Gittins policy for discovering compact, high-value configuration sets. For each configuration, it maintains one vector posterior over both objectives, shared across radial scalarizations, and uses direction-specific indices to decide whether another batch is worth its monetary search cost. AgentOpt, our client-side implementation, measures role-attributed outcomes, token usage, monetary cost, and latency and recommends Pareto configurations. We evaluate radial-Gittins on HotpotQA planner--solver and MathQA answer--critic workflows using exhaustive $9 \times 9=81$ assignment matrices as replay data and ground truth. Across 20 matched replay seeds, it stops after spending, on average, only $5.9\%$ and $5.7\%$ of exhaustive-evaluation API cost, respectively, while returning compact sets with low hypervolume regret and generational distance.
Chat is not available.
Successful Page Load