EvoTrace: Execution-Verified Visual Programming
Abstract
Visual programming agents produce execution histories that final-answer evaluation discards. We introduce EvoTrace, a framework for test-time improvement through verified transitions linking programs, renders, critiques, actions, and execution outcomes. Only executable programs become committed states; the generator retains their history, while the verifier receives a separate context. On all 280 BlenderGym tasks with a shared frozen backbone and evaluator, removing trajectory context yields paired error differences of 35.1% on N-CLIP and 45.3% on PL in favor of the full system, and increases repeated execution errors from 17.6% to 48.6%. Sharing generator and verifier histories degrades PL, whereas verifier persistence and local-edit constraints show no resolved benefit. With four candidates per round, EvoTrace completes all tasks and reduces both metrics by 32–38% against VIGA on jointly completed tasks. At matched candidate breadth, the advantage remains established only on N-CLIP. A replicated, split-controlled skill-consolidation experiment finds no statistically resolved cross-task transfer. These results support verified trajectory context for within-task refinement while delimiting claims about self-evolution.