VisionCreator-R1: A Reflection-Enhanced Native Visual-Generation Agentic Model
Abstract
Visual content generation has moved from single-image to multi-image workflows, yet existing agents are largely plan-driven and lack a systematic reflection mechanism for correcting mid-trajectory visual errors. To address this gap, we propose VisionCreator-R1, a native visual generation agent with explicit reflection, paired with a Reflection-Plan Co-Optimization (RPCO) training methodology. Through extensive experiments and trajectory-level analysis, we uncover a reflection-plan optimization asymmetry in reinforcement learning (RL): planning can be reliably optimized via plan rewards, while reflection learning is held back by noisy credit assignment. Motivated by this finding, we further abstract a general Decouple-then-Fuse paradigm for co-optimizing capabilities with asymmetric reward variance, of which RPCO is the visual-generation instantiation. Our RPCO first trains on the self-constructed VCR-SFT dataset, which covers reflection-strong single-image trajectories and planning-strong multi-image trajectories, and then co-optimizes on the VCR-RL dataset via RL. This yields our unified VisionCreator-R1 agent, which beats Gemini2.5Pro on existing benchmarks and on our VCR-Bench covering single-image and multi-image tasks.