VisionCreator-S1: Evolving Visual-Generation Agents via Skill-GRPO Optimization
Abstract
Native visual-generation agents now handle long-horizon, multi-image workflows, but face two practical bottlenecks during reinforcement learning (RL) training. First, the inability to accumulate experience across rollouts prevents the agent from leveraging past successes and failures, slowing the evolution of its agent behaviors. Second, the reliance on black-box visual-generation tools introduces tool-induced reward noise, as the policy is unfairly penalized for failures rooted in a tool's inherent capability boundary, leading to unstable and inefficient training. To address these challenges, we propose VisionCreator-S1, an agent that casts skill-augmented RL as a nested optimization problem through two coupled innovations. First, Dual-Skills co-evolves an agent-skill set to reuse self-reflection patterns from historical rollouts, and a tool-skill set that dynamically maps each tool's capability boundary through world-reflection. Second, Skill-GRPO treats discrete text skills as learnable variables. To overcome the non-differentiability of text skill, it gracefully alternates between numerical gradient descent for the policy weights and textual gradient descent for the discrete skill sets. By contrasting high- and low-advantage trajectories, it approximates a direction-consistent textual gradient in the agent's behavior feature space. Extensive experiments on VCR-Bench, GEdit-Bench, and MultiBanana demonstrate that VisionCreator-S1 consistently outperforms the strongest skill-free baseline, VisionCreator-R1. Ablation studies further confirm that this aligned co-evolution significantly improves both training stability and the agent's long-term decision-making capabilities.