VGI-BENCH: Probing Visual Intelligence in Video Generation Models

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为评估视频生成模型的视觉推理能力,研究引入了VGI-bench基准测试,包含27项任务,用以挑战并分析现有模型在零样本视觉推理中的表现与局限。
📝 Abstract
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce VGI-bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models. Our evaluations show that current generative systems can solve a subset of visually grounded reasoning tasks, but remain far from reliable, with even the strongest model, Seedance~2.0, achieving only 51.0% under our evaluation criteria. Our analysis further explore the output failure modes, input condition sensitivity, performance transfer boundary from synthetic fine-tuning, and internal denoising perspective revealing limited self-correction, where later steps mainly refine early hypotheses rather than correct reasoning errors. We hope VGI-bench will help stimulate the development of next-generation video generation models. We will release our code and data.
Problem

Research questions and friction points this paper is trying to address.

video generation models
visual reasoning
evaluation benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

VGI-bench
visual reasoning
video generation models
benchmarking
performance evaluation
X
Xuan He
University of Illinois Urbana Champaign
Cong Wei
Cong Wei
University of Waterloo
ReasoningDiffusionEfficiency
Y
Yuhao Cheng
University of Illinois Urbana Champaign
L
Linrui Ma
Tsinghua University, Massachusetts Institute of Technology
Y
Yuxuan Zhang
University of British Columbia, Vector Institute, University of Illinois Urbana Champaign
Z
Zuojun Li
Tsinghua University
Y
Yuhao Wen
Tsinghua University
Zeyi Liu
Zeyi Liu
Tsinghua University
Safety-guaranteed ControlSafety AssessmentFault DiagnosisOnilne Learning
Yuren Hao
Yuren Hao
Student, University of Illinois at Urbana-Chanpaign
NLPBio-inspired AI
S
Songcheng Cai
University of Waterloo, Vector Institute
Keming Wu
Keming Wu
Ph.D. Student, Tsinghua University
Computer VisionVision Language ModelsGenerative AI
Penghui Du
Penghui Du
Southern University of Science and Technology, Undergraduate
NeuroscienceMachine LearningfMRI imaging
Kai Zou
Kai Zou
Founder CEO, ProtagoLabs, NetMind.ai and AGI odyssey
Artificial General Intelligence
R
Rui Yang
University of Illinois Urbana Champaign
Chenkai Sun
Chenkai Sun
University of Illinois at Urbana-Champaign
Natural Language ProcessingDeep LearningArtificial Intelligence
Ke Yang
Ke Yang
University of Illinois at Urbana-Champaign
Natural Language ProcessingMachine Learning
Ping Nie
Ping Nie
Waterloo University
Natural Language ProcessingInformation RetrievalRecommendation SystemsTime Series Forecasting
K
Kelsey R Allen
University of British Columbia, Vector Institute
C
Chenglong Wang
Microsoft Research
Michel Galley
Michel Galley
Sr. Principal Research Manager at Microsoft
natural language processingdeep learningmachine learning
Jianfeng Gao
Jianfeng Gao
Microsoft Research, Redmond
natural language processinginformation retrievalmachine learningdeep learning
ChengXiang Zhai
ChengXiang Zhai
University of Illinois at Urbana-Champaign
Intelligent Information SystemsIntelligent AgentsFoundation ModelsHealthcareEducation