When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建包含复杂组成因素的图像缺陷数据集CO-AID,利用人类评估来识别AI生成图像中的缺陷,并训练深度模型以预测和优化这些缺陷。
📝 Abstract
*Chulin Zhao and Ruoqi Hu contributed equally to this work. State-of-the-art text-to-image (T2I) models exhibit pronounced and systematic defects when prompts involve intricate compositional factors such as multiple entities and multiple attributes. In this paper, we investigate how humans identify such defects. Specifically, we manually select 651 reference images from the four categories of people, hand, object, and scene that exhibit complex compositional characteristics, from which prompts emphasizing compositional factors are derived by manually editing ChatGPT-generated prompts. We then feed the prompts into three selected T2I models to generate AI images and conduct a comprehensive subjective study to identify their defects. For each image, 29 participants provide multi-label assessments specifying defect types and locations. The study yields the compositional AI-generated image defect (CO-AID) dataset, including reference images, prompts, AI-generated images, and information on defect locations and types. Experimental results show that training a deep model on CO-AID can both predict defects in AI-generated images and optimize AI image generation, demonstrating its usability and effectiveness. The database and supplementary materials are available at: https://github.com/Future-IQA/CO-AID .
Problem

Research questions and friction points this paper is trying to address.

text-to-image
compositional factors
defects
Innovation

Methods, ideas, or system contributions that make the work stand out.

compositional AI-generated image defect (CO-AID) dataset
multi-label assessments
defect prediction and optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruoqi Hu
Dundee International Institute of Central South University, Central South University, China
C
Chulin Zhao
Dundee International Institute of Central South University, Central South University, China
J
Jiashuo Chang
Dundee International Institute of Central South University, Central South University, China
R
Ramon Ruiz-Dolz
Faculty of Science, Engineering and Business, University of Dundee, United Kingdom
Hanhe Lin
Hanhe Lin
SSEN@University of Dundee
visual quality assessmentmedical image analysiscomputer vision