🤖 AI Summary
This work addresses the limited cultural grounding of current vision-language models, which—due to their reliance on English-centric training data—struggle to capture multimodal associations within non-English cultural contexts such as Polish. To bridge this gap, the study introduces PoVisLE, the first deep evaluation benchmark specifically designed for a single non-English culture. PoVisLE comprises 1,117 images paired with 2,366 expert-annotated visual question-answer pairs, curated under an embodied evaluation paradigm that emphasizes the interplay between language and vision within concrete cultural settings. Moving beyond superficial, template-based recognition tasks, this benchmark offers a high-quality, controllable, and challenging resource to advance the cultural adaptation and evaluation of vision-language models beyond English-dominated frameworks.
📝 Abstract
Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predominantly trained on English-centric data, which limits their ability to handle culturally grounded visual understanding and leads to failures in interpreting region-specific meanings, symbolic content, and context-dependent visual cues. Existing benchmarks for cultural competence are often template-driven and focused on surface-level recognition, making them insufficient for evaluating deeper linguistic and pragmatic understanding in culturally situated settings. We introduce PoVisLE, a monocultural vision-language benchmark for Polish designed to evaluate culturally grounded multimodal understanding under a grounded evaluation paradigm, where language is interpreted in interaction with visual context. The dataset contains 1,117 images and 2,366 manually annotated VQA pairs. Overall, our dataset provides a controlled and challenging resource for assessing culturally grounded vision-language understanding beyond surface-level recognition.