Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models

πŸ“… 2026-08-12
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing evaluation methods treat visual generation and understanding as disjoint capabilities, failing to holistically assess the system-level performance of unified multimodal models. This work proposes a Self-Generated Understanding (SGU) framework, introducing a novel annotation-free semantic closed-loop evaluation paradigm: a model first describes an image, then reconstructs visual content from its own generated text, and finally performs zero-shot reasoning on the reconstructed output. SGU uniquely integrates generation and understanding into a single assessment pipeline, leveraging a self-feedback mechanism to uncover latent deficiencies that manifest specifically within the model’s self-generated context. Experiments reveal that even high-performing models exhibit substantially degraded reasoning capabilities under SGU, exposing systemic limitations invisible to conventional isolated evaluations.
πŸ“ Abstract
As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs). In this work, we propose Self-Generative-Understanding (SGU), a novel, annotation-free evaluation framework that probes the integrated capabilities of unified models through a semantic closed-loop challenge. Without requiring new annotations, SGU leverages the dual understanding-and-generation abilities of UMMs by asking them to first perceive an image and produce a textual description, subsequently reconstruct a visual context based on that description, and finally perform reasoning over the self-generated output. This pipeline provides a zero-cost testbed that yields an integrated performance score specifically tailored for evaluating UMMs as unified systems. Extensive experiments show that even high-performing UMMs often struggle to reason over their own generated contexts, revealing limitations that are not captured by separate evaluations of understanding or generation alone. Our work provides a complementary holistic evaluation framework and offers a foundation for benchmarking the development of next-generation unified multimodal models.
Problem

Research questions and friction points this paper is trying to address.

unified multimodal models
holistic evaluation
vision-language models
semantic closed-loop
integrated capabilities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Generative-Understanding
semantic closed-loop
unified multimodal models
holistic evaluation
annotation-free
πŸ”Ž Similar Papers
2024-02-22Computer Speech & LanguageCitations: 2
πŸ’Ό Related Jobs
No related jobs found.
H
Hao Zhang
Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences
J
Jiaxin Qi
Computer Network Information Center, Chinese Academy of Sciences
Zhijiang Tang
Zhijiang Tang
Postgraduate student at University of Chinese Academy of Sciences
Deep LearningAI for ScienceTime Series Analyze
Jianqiang Huang
Jianqiang Huang
Nanyang Technological University, Chinese Academy of Sciences
Compter VisionMachine LearningCasuality