MMArt A Multi-Perspective Multimodal Dataset for Visual Art Understanding

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing vision-language models for art understanding are largely confined to superficial descriptions and lack deep, multidimensional interpretation across narrative, formal, emotional, and historical dimensions. Moreover, no unified multimodal dataset with multi-perspective annotations currently exists. To address this gap, this work introduces MMArt, a large-scale multimodal art dataset comprising 74,234 WikiArt paintings, each annotated with four distinct perspective-specific descriptions—narrative, formal, emotional, and historical—and a unified integrative caption. The study further proposes a novel four-dimensional collaborative annotation framework. Through generative reconstruction, discriminative retrieval, and ablation studies, the authors demonstrate the irreplaceable role of each perspective: narrative descriptions achieve the highest retrieval performance (R@1=44.0%), formal descriptions best preserve compositional style, and historical descriptions carry the strongest emotional signals. These findings underscore that single-perspective approaches are insufficient for diverse downstream tasks, thereby advancing art-focused AI toward deeper, multifaceted understanding.
📝 Abstract
Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface content but struggle with formal analysis, grounded historical interpretation, or affective characterization. We argue this is not only a model but also a dataset limitation. Existing art datasets are single perspective resources, where no dataset provides narrative, formal, emotional, and historical perspectives simultaneously for the same artworks. We introduce MMArt, a large-scale dataset of 74,234 WikiArt paintings, each annotated with four independently annotated perspectives plus a harmonized unified caption, produced by specialized vision-language models or human annotation and validated through complementary quality evaluations. Two complementarity analyses establish that perspectives encode genuinely distinct information. A generative analysis shows that formal analysis descriptions best preserve compositional style, and historical descriptions carry strong affective signal in reconstructed images. A discriminative retrieval analysis reveals task-asymmetry: narrative descriptions drive retrieval (R@1 = 44.0%), while formal descriptions, strongest for reconstruction, are nearly nondiscriminative at retrieval scale (R@1 = 7.8%). Leave-one-out analysis further confirms that historical descriptions are the least replaceable perspective across both tasks. Together, the two analyses establish that no single perspective suffices for all tasks, directly motivating MMArt multi-perspective design. The dataset, code, and additional information are available at https://shuaiwang97.github.io/MMArt/.
Problem

Research questions and friction points this paper is trying to address.

visual art understanding
multimodal dataset
multi-perspective annotation
vision-language models
art interpretation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-perspective
multimodal dataset
visual art understanding
complementary analysis
formal analysis
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Shuai Wang
Shuai Wang
University of Amsterdam
Multimodal RetrievalGenerative AITopological Deep Learning
W
Wangyuan Ding
University of Amsterdam, Amsterdam, The Netherlands
Yixian Shen
Yixian Shen
University of Amsterdam
Efficient DNNComputer ArchitectureSystem Optimization
J
Jia-Hong Huang
University of Amsterdam, Amsterdam, The Netherlands; Amazon AGI, United States of America
Stevan Rudinac
Stevan Rudinac
Associate Professor, University of Amsterdam
multimediacomputer visioninformation retrievalmachine learning
M
Monika Kackovic
University of Amsterdam, Amsterdam, The Netherlands
Nachoem Wijnberg
Nachoem Wijnberg
Unknown affiliation
M
Marcel Worring
University of Amsterdam, Amsterdam, The Netherlands