CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing image editing benchmarks, which suffer from narrow evaluation dimensions and inadequate assessment of complex multi-image reasoning in real-world deployment scenarios. To overcome these challenges, we propose CPI-Bench, a novel benchmark that introduces pioneering multi-image editing evaluation metrics. It establishes a hierarchical framework comprising general, practical, and intelligent subsets, validated through comprehensive human preference alignment analysis. Experimental results demonstrate that CPI-Bench significantly enhances model performance discrimination and exhibits high correlation with human judgments. Consequently, this benchmark accurately quantifies real-world editing and reasoning capabilities, providing reliable guidance for future model optimization and development in practical applications.
📝 Abstract
With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical andIntelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and pioneers the inclusion of multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating it faithfully captures the preferences and perceptual judgments of human evaluators, serving as a robust proxy for real-world user experience.
Problem

Research questions and friction points this paper is trying to address.

Image Editing Evaluation
Benchmark Limitations
Multi-image Editing
Reasoning-based Editing
Real-world Deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-image Editing Evaluation
Reasoning-based Editing
Real-world Benchmark
Human Alignment
Performance Differentiation
Q
Qinye Zhou
Alibaba Group
J
Jun Zheng
Alibaba Group
Y
Yongchao Du
Alibaba Group
Y
Yuan Wang
Alibaba Group
Z
Zhengrui Chen
Alibaba Group
Zuan Gao
Zuan Gao
University of Science and Technology
GenAIAIGCOCR
Taihang Hu
Taihang Hu
Nankai University
Deep LearningComputer VisionGenerative models
C
Chao Lin
Alibaba Group
Y
Yefeng Shen
Alibaba Group
X
Xingjian Wang
Alibaba Group
Z
Zhao Wang
Alibaba Group
Z
Zhengtao Wu
Alibaba Group
Xiaoli Xu
Xiaoli Xu
Southeast University, China
Wireless communicationnetwork codingchannel coding
Z
Zhengze Xu
Alibaba Group
H
Hao Yan
Alibaba Group
D
Denghui Yang
Alibaba Group
Y
Yuhang Yu
Alibaba Group
Huayu Zhang
Huayu Zhang
Senior Engineer, Huawei Technologies Co., Ltd
Distributed SystemNetwork ScienceMachine LearningOptimizationGraph Theory
M
Mingzhou Zhang
Alibaba Group
Mengting Chen
Mengting Chen
Alibaba Group
Generative ModelingComputer Vision