Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited cultural grounding of current vision-language models, which—due to their reliance on English-centric training data—struggle to capture multimodal associations within non-English cultural contexts such as Polish. To bridge this gap, the study introduces PoVisLE, the first deep evaluation benchmark specifically designed for a single non-English culture. PoVisLE comprises 1,117 images paired with 2,366 expert-annotated visual question-answer pairs, curated under an embodied evaluation paradigm that emphasizes the interplay between language and vision within concrete cultural settings. Moving beyond superficial, template-based recognition tasks, this benchmark offers a high-quality, controllable, and challenging resource to advance the cultural adaptation and evaluation of vision-language models beyond English-dominated frameworks.
📝 Abstract
Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predominantly trained on English-centric data, which limits their ability to handle culturally grounded visual understanding and leads to failures in interpreting region-specific meanings, symbolic content, and context-dependent visual cues. Existing benchmarks for cultural competence are often template-driven and focused on surface-level recognition, making them insufficient for evaluating deeper linguistic and pragmatic understanding in culturally situated settings. We introduce PoVisLE, a monocultural vision-language benchmark for Polish designed to evaluate culturally grounded multimodal understanding under a grounded evaluation paradigm, where language is interpreted in interaction with visual context. The dataset contains 1,117 images and 2,366 manually annotated VQA pairs. Overall, our dataset provides a controlled and challenging resource for assessing culturally grounded vision-language understanding beyond surface-level recognition.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
cultural competence
grounded understanding
multimodal evaluation
Polish
Innovation

Methods, ideas, or system contributions that make the work stand out.

culturally grounded vision-language understanding
monocultural benchmark
grounded evaluation paradigm
Polish VQA dataset
vision-language models
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
A
Anna Kołos
NASK National Research Institute, Warsaw, Poland
G
Grzegorz Statkiewicz
NASK National Research Institute, Warsaw, Poland
Karolina Seweryn
Karolina Seweryn
NASK - National Research Institute, Warsaw University of Technology
K
Katarzyna Kowol
NASK National Research Institute, Warsaw, Poland
K
Karolina Piosek
NASK National Research Institute, Warsaw, Poland
Wojciech Kusa
Wojciech Kusa
NASK National Research Institute
Natural Language ProcessingInformation RetrievalMachine LearningLLMs