HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
HUG-VIS通过提供统一的多模态基准,解决人类中心视觉智能任务中理解与生成的问题,涵盖情感识别、视频生成等四个任务。
📝 Abstract
Visual intelligence seeks to perceive, interpret, and synthesize the visual world and is central to modern computer vision. Human-centered visual intelligence is especially demanding because it studies people as expressive, socially situated subjects whose meaning is rarely conveyed by appearance alone. It couples vision with audio and language across four representative tasks: human emotion recognition, human video generation, human voice cloning, and human video matting. Yet existing resources remain task-specific, providing modalities and annotations for individual problems rather than a shared foundation coordinating understanding and generation. This limits multimodal signal use and broader research. We address this gap with HUG-VIS, a unified benchmark for Human-centered Understanding and Generation in Visual Intelligence. It contains 8,400 seated half-body videos of 30 professional actors, each performing the same 280 emotion-action-prompt assignments under a controlled Mandarin studio protocol, with synchronized video, audio, text, and alpha mattes. We evaluate diverse open- and closed-source models across the four tasks under a unified zero-shot protocol using automatic metrics, criterion-specific mean opinion scores, and multiple cross-task analyses. Results show that (i) linguistic content dominates current emotion recognition, while purely visual affect recognition is weakest; (ii) in video generation and voice cloning, automatic metrics and human judgment agree overall but differ in their top rankings, requiring joint reporting; (iii) boundary fidelity under motion is the main remaining obstacle for human matting; and (iv) task difficulty varies across emotions, models, and metrics, with notable cross-task correlations. The dataset and results are available at https://github.com/GML-MMGroup/HUG-VIS.
Problem

Research questions and friction points this paper is trying to address.

multimodal
visual intelligence
human-centered
understanding and generation
task-specific
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Benchmark
Human-centered Visual Intelligence
Zero-shot Protocol
Cross-task Analyses
🔎 Similar Papers
No similar papers found.
F
Fei Ma
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), Shenzhen, China.
Zebang Cheng
Zebang Cheng
Shenzhen University
AICVMLLMAffective Computing
Minghui Li
Minghui Li
Huazhong University of Science and Technology
AI Security
H
Hongbo Xu
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), Shenzhen, China.
Y
Yuyong Tan
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), Shenzhen, China.; Institute of Automation, Chinese Academy of Sciences, Beijing, China.
Y
Yihua Shao
Institute of Automation, Chinese Academy of Sciences, Beijing, China.
Hanling Wang
Hanling Wang
Tsinghua University
machine learningdeep learningdistributed system
Zhou Liu
Zhou Liu
China Southern Power Grid/ Shenzhen Power Supply Co., Ltd.
Renewable Power IntegrationSmart gridPower system protectionDigital substationAI technology
Y
Yuqing Gao
Tongji University, Shanghai, China.
D
Dong Wang
Tsinghua University, Beijing, China.
Long Ma
Long Ma
Dalian University of Technology
Computer VisionImage Processing
Laizhong Cui
Laizhong Cui
Shenzhen University
NetworkingEdge ComputingIoTBig Data,Machine Learning
Nicu Sebe
Nicu Sebe
University of Trento
computer visionmultimedia
Q
Qi Tian
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), Shenzhen, China.; Huawei, Shenzhen, China.