Understanding the Challenges in Iterative Generative Optimization with LLMs
Abstract
Generative optimization uses large language models (LLMs) to iteratively improve artifacts (like code, workflows or prompts) using execution feedback. It is a promising approach to building self-improving agents, yet in practice remains brittle: despite active research, only 9% of surveyed agents used any automated optimization. We argue that this brittleness arises because, to set up a learning loop, an engineer must make "hidden" design choices: What can the optimizer edit and what learning evidence is provided per update? We investigate three factors that affect most applications: the starting artifact, batching execution traces as experiences, and determining a credit horizon with truncated traces. Through case studies in MLAgentBench, Atari, and BigBench Extra Hard, we find that decisions around these factors can determine whether generative optimization succeeds and they are not made explicit in prior works. Different starting artifacts determine which solutions are reachable in MLAgentBench, truncated traces can still improve Atari agents, and larger minibatches do not monotonically improve generalization on BBEH. We conclude that the lack of a simple, universal learning-loop setup across domains is a major hurdle for productionization and adoption, and provide practical guidance for making these design choices.