Reasoning with Image Generation

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过将图像生成模型作为多模态大语言模型的灵活视觉推理机制,解决了需要直接操作视觉表示的任务问题。
📝 Abstract
Chain-of-thought reasoning has revolutionized natural language processing by enabling large language models (LLMs) to decompose problems into intermediate steps before answering. Yet confining reasoning to the textual domain presents limitations for tasks requiring direct manipulation of visual representations. Recent efforts augment multimodal LLMs with external visual expert tools such as depth estimation or object detection modules, but these remain fundamentally limited by their reliance on narrow, rigid operations that cannot flexibly generate or transform visual content. We propose ReImaGin, which leverages image generation models as a flexible visual reasoning mechanism for multimodal LLMs: unlike fixed-function tools, they accept natural language commands and can perform open-ended visual operations, like removing an occlusion or generating a floorplan from multiple disjoint views of a room. Across six diverse visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin consistently outperforms both text-only reasoning and specialist vision-tool baselines, with gains of up to 25\%, demonstrating the advantage of flexible, generative visual reasoning.
Problem

Research questions and friction points this paper is trying to address.

Chain-of-thought reasoning
Multimodal LLMs
Visual representation
Image generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

image generation models
flexible visual reasoning
natural language commands
open-ended visual operations
multimodal LLMs
🔎 Similar Papers
No similar papers found.
N
Nishad Singhi
Technical University of Darmstadt & hessian.AI
H
Hector Garcia Rodriguez
Technical University of Darmstadt & hessian.AI
A
Aditya Arora
Technical University of Darmstadt & hessian.AI
Marcus Rohrbach
Marcus Rohrbach
Professor for Multimodal Reliable AI, TU Darmstadt, Germany
Machine LearningComputer VisionAI
Anna Rohrbach
Anna Rohrbach
Professor, TU Darmstadt, Germany
Vision and LanguageArtificial IntelligenceMultimodal Grounded Learning