When Does Visual Generation Help Visual Understanding in Unified Multimodal Models?

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过VGAU-Diag框架评估视觉生成如何辅助统一多模态模型中的视觉理解,发现生成的视觉辅助在简单任务中有效,但在复杂推理时不可靠。
📝 Abstract
Unified multimodal models (UMMs) can perform both understanding and generation, raising a central question: can visual generation improve understanding? Existing evaluations provide mixed evidence, but confound task difficulty, reasoning paradigms, and the closed-loop interaction between generation and understanding. We introduce VGAU-Diag, a fine-grained evaluation framework for vision generation-assisted understanding. It stratifies samples by difficulty, enables unified evaluation of multiple reasoning paradigms, and uses Oracle-Assisted Reference Protocols. Our analysis shows that generated visual aids help on easier instances but become unreliable as reasoning complexity increases. Oracle-assisted diagnosis further reveals that the main bottleneck often lies on the visual-understanding side rather than the visual-generation side, as current UMMs struggle to leverage even faithful visual aids. We also show that effective visual generation should target visual-understanding bottlenecks rather than add more reasoning steps, and identify a three-stage transition from task-irrelevant noise, to misleading plausible guidance, and finally to useful assistance. These findings would be useful to guide the development of better UMMs.The code is available at https://github.com/zyb1029/VGAU-Diag.
Problem

Research questions and friction points this paper is trying to address.

Unified Multimodal Models
Visual Generation
Visual Understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

VGAU-Diag
visual generation-assisted understanding
unified multimodal models
reasoning complexity
visual-understanding bottlenecks