Understanding Diversity in LLM-Based Algorithm Design
Abstract
Diversity is a central design principle in LLM-based algorithm design, and prominent methods include mechanisms to maintain or promote diversity across the generation, evaluation, and selection loop that produces heuristic programs. Yet what diversity actually contributes to performance remains unclear. We address this gap with DivEval, a unified empirical framework that measures diversity at the code, logic, behavior, and performance-profile levels across search methods, language-model backbones, and tasks. The macro lens relates set-level diversity to held-out performance and finds near-optimal regions whose location and width depend on both the task and the measurement level, making diversity a quantity to calibrate rather than to maximize. The temporal lens tracks diversity during search and shows that the generator keeps producing varied candidates while the retained population contracts into a high-retention, low-novelty state. The mechanistic lens finds that the levels do not collapse into one: code and logic form one tightly coupled block, behavior and performance profile another, with only a weak bridge between them, so source-level novelty is a poor proxy for functional novelty. Together, the results characterize diversity in LLM-based algorithm design as a task-dependent, temporally evolving, multi-level property that should be measured and managed accordingly.