PosterDuet: Co-Evolving Design Generation and Reward Optimization for Product Poster Synthesis
Abstract
Automatic e-commerce poster generation requires the joint optimization of product fidelity, visual composition, text rendering, and commercial appeal. Existing methods typically address this task through staged pipelines involving product segmentation, layout prediction, glyph rasterization, and conditional image synthesis. While effective in constrained settings, such decompositions rely on rigid preprocessing assumptions and weaken the mutual adaptation among product appearance, typography, and scene composition, especially for hand-held, worn, or context-dependent products. We reformulate this problem as \emph{holistic product-aware poster editing}: given a raw product image and structured metadata, the goal is to generate a complete poster directly, without segmentation masks, glyph control maps, or predefined layout boxes. To this end, we propose PosterDuet, a closed-loop framework that integrates a vision-language model (VLM), an image editing model, and reward-based optimization. The VLM first generates a holistic design prompt that specifies background style, layout arrangement, promotional copy, and typographic intent from the raw image and metadata. Conditioned on both the prompt and the original image, the image editor synthesizes the final poster in a unified editing process. To optimize both commercial effectiveness and visual quality, we introduce a mixed-reward learning framework that combines a CTR-oriented reward with a generative holistic quality reward, and use GRPO to optimize the prompt-generating VLM. We further improve textual accuracy via OCR-reward-based reinforcement tuning of the image editor, and exploit natural-language critiques from the generative reward model for iterative prompt refinement. Extensive experiments show that PosterDuet generates more coherent, text-faithful, and commercially effective posters than prior pipelined approaches.