VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage

📅 2025-10-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing VQA benchmarks primarily emphasize superficial visual attributes, failing to adequately assess models’ deep semantic understanding in art and cultural heritage domains—e.g., symbolic meaning, narrative structure, and cultural context. To address this gap, we introduce ArtVQA, the first large-scale, domain-specific visual question answering benchmark for art and cultural heritage. Our method employs a multi-agent collaborative generation framework, where a domain-expert agent orchestrates question design and validation to ensure linguistic diversity, semantic richness, and comprehensive coverage of multi-dimensional visual understanding—including fine-grained object recognition, relational reasoning, and cultural metaphor interpretation. Systematic evaluation across 14 state-of-the-art multimodal large language models reveals pervasive deficits in deep reasoning, particularly in counting, cross-modal alignment, and cultural modeling. Notably, our study uncovers, for the first time, a substantial performance gap between open-source and closed-source models on this benchmark.

Technology Category

Application Category

📝 Abstract
Multimodal Large Language Models (MLLMs) have demonstrated significant capabilities in joint visual and linguistic tasks. However, existing Visual Question Answering (VQA) benchmarks often fail to evaluate deep semantic understanding, particularly in complex domains like visual art analysis. Confined to simple syntactic structures and surface-level attributes, these questions fail to capture the diversity and depth of human visual inquiry. This limitation incentivizes models to exploit statistical shortcuts rather than engage in visual reasoning. To address this gap, we introduce VQArt-Bench, a new, large-scale VQA benchmark for the cultural heritage domain. This benchmark is constructed using a novel multi-agent pipeline where specialized agents collaborate to generate nuanced, validated, and linguistically diverse questions. The resulting benchmark is structured along relevant visual understanding dimensions that probe a model's ability to interpret symbolic meaning, narratives, and complex visual relationships. Our evaluation of 14 state-of-the-art MLLMs on this benchmark reveals significant limitations in current models, including a surprising weakness in simple counting tasks and a clear performance gap between proprietary and open-source models.
Problem

Research questions and friction points this paper is trying to address.

Evaluating deep semantic understanding in visual art analysis
Overcoming limitations of simple syntactic VQA benchmarks
Assessing symbolic meaning interpretation in cultural heritage images
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-agent pipeline generates nuanced art questions
Benchmark probes symbolic meaning and visual relationships
Specialized agents create validated linguistically diverse queries
🔎 Similar Papers
💼 Related Jobs
No related jobs found.