Build, Test, Success: Verification in the Loop Turns LLM Agents into Brick Builders
Abstract
Can a language model build a physical structure? We study this question with LEGO-style bricks, a domain where the parts are fixed, expert data is plentiful, and every claim a build makes—parts connect, nothing overlaps, the structure stands—can be checked by machine. We propose multi-step iterative brick generation with error recovery: a general frontier model builds inside an interactive loop where each placement is expressed in a connection language and checked as it is made, and a validator reports which parts are wrong and how, not just that the build failed. On a suite of thirteen tasks with exact automatic grading, direct one-shot generation passes two tasks—the two most repetitive shapes—while the same model inside the loop passes twelve of the thirteen, including a 470-part Taj Mahal. We also perform a five-condition ablation to determine which component is responsible: the connection language provides physical validity, a single verification-and-repair pass at the end provides the rest, and letting the agent invoke checks mid-build adds nothing further. The environment, the validator with its measured error, the tasks, and every recorded build session will be released on GitHub upon acceptance, allowing full reproduction of these results.