Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model

📅 2026-07-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current AI-generated human-centric videos often suffer from quality artifacts and semantic inconsistencies, and lack comprehensive multidimensional evaluation benchmarks. To address this gap, this work introduces HVEval+, the first large-scale human-annotated dataset encompassing three critical dimensions: spatial quality, temporal coherence, and text-video alignment. Furthermore, we propose MoE-Rater, a unified multitask evaluation model that integrates Mixture-of-Projection Experts (MoPE) and Mixture-of-LoRA Experts (MoLE), trained via a three-stage strategy. Evaluated on both HVEval+ and Human-AGVQA, MoE-Rater significantly outperforms existing methods, offering a reliable and holistic assessment tool to advance text-to-video generation models.
📝 Abstract
AI-generated human-centric videos play a crucial role in a wide range of modern applications. However, they often suffer from quality issues and semantic mismatches, underscoring the importance of effective quality assessment for such videos. To this end, we extend our previous dataset HVEval with pairwise preference annotations, resulting in HVEval+, the largest holistic quality assessment dataset for AI-generated human-centric videos, which comprises 1k prompts based on a comprehensive taxonomy, 20k videos generated by 24 text-to-video (T2V) models, and extensive human annotations, including 60k mean opinion scores (MOSs) and 60k preference pairs across 3 dimensions (i.e., spatial quality, temporal quality, and text-video correspondence), as well as 20k category-specific question-answer (Q&A) pairs. Along with the HVEval+ dataset, we further propose MoE-Rater, a Mixture-of-Experts (MoE)-inspired and multimodal large language model (MLLM)-based all-in-one method that supports multi-dimensional quality rating, multi-dimensional pairwise comparison, and category-specific question answering within a single model. Specifically, we introduce Mixture of Projector Experts (MoPE) and Mixture of LoRA Experts (MoLE), together with a three-stage training strategy consisting of task-aware pre-training, task-specific adaptation, and adaptive routing optimization, to effectively unify multiple tasks, resulting in superior performance on both HVEval+ and Human-AGVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the HVEval+ dataset and the MoE-Rater method in advancing AI-generated video quality assessment and further facilitating the evaluation and optimization of T2V models.
Problem

Research questions and friction points this paper is trying to address.

AI-generated videos
quality assessment
human-centric videos
text-to-video models
semantic mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
multimodal large language model
multi-dimensional quality assessment
text-to-video evaluation
HVEval+
S
Sijing Wu
Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai, China
Y
Yunhao Li
Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai, China
Huiyu Duan
Huiyu Duan
Shanghai Jiao Tong University
Multimedia Signal Processing
Yucheng Zhu
Yucheng Zhu
Shanghai Jiaotong University
Multimedia Signal Processing
X
Xiongkuo Min
Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai, China
Patrick Le Callet
Patrick Le Callet
Prof. Universite de Nantes, LS2N, Polytech Nantes - Institut Universitaire de France (IUF)
cognitive computing for MMQoEhuman perception and applications in ICT
Guangtao Zhai
Guangtao Zhai
Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI EvaluationDisplays