Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对视觉生成模型缺乏高层次推理能力的问题,提出RIG-BENCH基准,评估模型在四个认知领域中的推理驱动图像生成能力。
📝 Abstract
Recent advancements in unified generative models (UGMs) and world simulators have achieved unprecedented results in visual perception and synthesis. However, these models primarily rely on surface-level event alignment, leaving the capacity for high-level visual reasoning underexplored. True visual generative intelligence demands "Reasoning-to-Generation", an ability to infer latent rules from visual inputs and manifest solutions through precise, logically constrained visual outcomes. We introduce RIG-BENCH, a novel comprehensive benchmark that systematically evaluates Reasoning-driven Image Generation (RIG) across four cognitively demanding domains: Concept-based, Transformation-based, Pattern & Structure, and Scenario-based. Featuring 2000 curated samples, RIG-BENCH serves as a rigorous stress test for RIG. Our extensive evaluations of state-of-the-art UGMs and image/video generation models reveal a significant reasoning-generation gap, wherein models frequently produce locally plausible but globally illogical outputs. RIG-BENCH provides a vital diagnostic framework to guide the development of next-generation, logically grounded UGMs and world simulators.
Problem

Research questions and friction points this paper is trying to address.

Reasoning-driven Image Generation
visual reasoning
unified generative models
Innovation

Methods, ideas, or system contributions that make the work stand out.

RIG-BENCH
Reasoning-driven Image Generation
visual reasoning
unified generative models
world simulators