Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation
Abstract
Recent advancements in world models and unified generative models (UGMs) have achieved unprecedented results in visual perception and synthesis. However, these models primarily rely on surface-level alignment, leaving the capacity for high-level visual reasoning underexplored. True visual generative intelligence demands "Reasoning-to-Generation", an ability to infer latent rules from visual inputs and manifest solutions through precise, logically constrained visual outcomes. We introduce RIG-Bench, a novel comprehensive benchmark that systematically evaluates Reasoning-driven Image Generation (RIG) across four cognitively demanding domains: Concept-based, Transformation-based, Pattern & Structure, and Scenario-based. Featuring 2000 curated samples, RIG-Bench serves as a rigorous stress test for RIG. Our extensive evaluations of state-of-the-art UGMs and image/video generation models reveal a significant reasoning-generation gap, wherein models frequently produce locally plausible but globally illogical outputs. RIG-Bench provides a vital diagnostic framework to guide the development of next-generation, logically grounded world models.