Born, Not Trained: Zero-Training Programs Out-Explore Reinforcement Learning for Game Testing at Equal Budget
Andy Ye
Abstract
Testing a game or simulation means driving it into rare states - the trigger for a planted bug or a level's own win condition - and reaching them within the small interaction budget a QA pass can afford. The standard approach trains a reinforcement-learning explorer for each game, which is expensive, needs per-game tuning, and rarely transfers. We propose a different route. Rather than train a tester, an autonomous LLM-evolution loop writes one as a short program. The pipeline is an iterative evolutionary search, not a single prompt: over many generations, a proposal LLM writes and mutates candidate programs, and a multi-engine fitness keeps the ones that reach the most target states at a fixed budget, re-verifying every reported hit from a clean reset. Each candidate is a compact procedure over a save/restore archive of simulator states, saving promising states and returning to them to explore further. The winning program needs no training: it explores through an adaptive novelty signal and a small bandit over exploration moves that emerged from the search, and it is written once and reused. A single fresh-per-task prompt does not suffice; the strength comes from evolving one program against a diverse multi-engine fitness and transferring it. At an equal interaction budget, this single written program is the best method on four sparse-reward game-testing tasks spanning three engines, reaching on average 91% of the targets against 54% for tuned PPO, and less for every other baseline. It matches what trained RL needs roughly $16\times$ more interaction to reach, and it transfers to engines the search never saw. However, the advantage is bounded. It ties two saturated tasks a trivial baseline also solves, loses the one reward-dense task to PPO, and is overtaken only when PPO is given several times more interaction, so writing and training a tester are complementary regimes. The Kinetix planted-defect benchmark contributes the replay-from-reset verification backbone, and that same replay verifier caught the loop inflating its own score through a race in its evaluation harness, a self-check we expect autonomous research will increasingly need.
Chat is not available.
Successful Page Load