🤖 AI Summary
Existing VQA benchmarks primarily emphasize superficial visual attributes, failing to adequately assess models’ deep semantic understanding in art and cultural heritage domains—e.g., symbolic meaning, narrative structure, and cultural context. To address this gap, we introduce ArtVQA, the first large-scale, domain-specific visual question answering benchmark for art and cultural heritage. Our method employs a multi-agent collaborative generation framework, where a domain-expert agent orchestrates question design and validation to ensure linguistic diversity, semantic richness, and comprehensive coverage of multi-dimensional visual understanding—including fine-grained object recognition, relational reasoning, and cultural metaphor interpretation. Systematic evaluation across 14 state-of-the-art multimodal large language models reveals pervasive deficits in deep reasoning, particularly in counting, cross-modal alignment, and cultural modeling. Notably, our study uncovers, for the first time, a substantial performance gap between open-source and closed-source models on this benchmark.
📝 Abstract
Multimodal Large Language Models (MLLMs) have demonstrated significant capabilities in joint visual and linguistic tasks. However, existing Visual Question Answering (VQA) benchmarks often fail to evaluate deep semantic understanding, particularly in complex domains like visual art analysis. Confined to simple syntactic structures and surface-level attributes, these questions fail to capture the diversity and depth of human visual inquiry. This limitation incentivizes models to exploit statistical shortcuts rather than engage in visual reasoning. To address this gap, we introduce VQArt-Bench, a new, large-scale VQA benchmark for the cultural heritage domain. This benchmark is constructed using a novel multi-agent pipeline where specialized agents collaborate to generate nuanced, validated, and linguistically diverse questions. The resulting benchmark is structured along relevant visual understanding dimensions that probe a model's ability to interpret symbolic meaning, narratives, and complex visual relationships. Our evaluation of 14 state-of-the-art MLLMs on this benchmark reveals significant limitations in current models, including a surprising weakness in simple counting tasks and a clear performance gap between proprietary and open-source models.