Terrarium: A Systematic Evaluation of LLM-Driven Optimization Systems
Abstract
LLM-driven optimization systems have shown success on prompts, agent scaffolds, and even arbitrary programs. As optimization systems report on their own benchmarks, budgets, and configurations, randomness in every aspect of these optimization systems makes their performance difficult to understand or compare. We propose Terrarium, a novel evaluation framework for controlled evaluations of LLM-driven optimization systems. With Terrarium, we are the first to systematically compare three types of optimization systems: LLM-based optimizers, agent-based optimizers, and fully autonomous agents. To our surprise, none of the studied optimizers is a clear winner across all the tasks. Based on findings from Terrarium, we propose a new optimizer pipeline called OMNI. Under the same budget, OMNI performs 14% to 41% better than any existing optimizer for program optimization tasks.