How Culturally Aware are Vision-Language Models?

📅 2024-05-24

🏛️ arXiv.org

📈 Citations: 0

✨ Influential: 0

career value

182K/year

🤖 AI Summary

This paper addresses the pervasive cultural inaccuracy of mainstream vision-language models (VLMs) in describing folkloric imagery—such as mythological scenes, traditional dances, and cultural symbols. To tackle this, we propose the first quantifiable, culture-aware evaluation framework. Methodologically, we introduce the Culture-Awareness Score (CAS), a novel metric for assessing cultural fidelity; curate MOSAIC-1.5k, an open-source dataset comprising 1,500 cross-cultural folkloric images with fine-grained cultural annotations; and adopt a hybrid evaluation paradigm integrating multi-model benchmarking, semantic alignment analysis, and expert human evaluation. Experiments reveal significant cultural blind spots across GPT-4V, Gemini Pro Vision, LLaVA, and OpenFlamingo; CAS effectively discriminates their relative cultural sensitivity. Both MOSAIC-1.5k and CAS are publicly released, establishing foundational benchmarks and evaluation standards for culturally adaptive AI systems.

Technology Category

Application Category

📝 Abstract

An image is often considered worth a thousand words, and certain images can tell rich and insightful stories. Can these stories be told via image captioning? Images from folklore genres, such as mythology, folk dance, cultural signs, and symbols, are vital to every culture. Our research compares the performance of four popular vision-language models (GPT-4V, Gemini Pro Vision, LLaVA, and OpenFlamingo) in identifying culturally specific information in such images and creating accurate and culturally sensitive image captions. We also propose a new evaluation metric, the Cultural Awareness Score (CAS), which measures the degree of cultural awareness in image captions. We provide a dataset MOSAIC-1.5k labeled with ground truth for images containing cultural background and context and a labeled dataset with assigned Cultural Awareness Scores that can be used with unseen data. Creating culturally appropriate image captions is valuable for scientific research and can be beneficial for many practical applications. We envision our work will promote a deeper integration of cultural sensitivity in AI applications worldwide. By making the dataset and Cultural Awareness Score available to the public, we aim to facilitate further research in this area, encouraging the development of more culturally aware AI systems that respect and celebrate global diversity.

Problem

Research questions and friction points this paper is trying to address.

Evaluating cultural awareness in vision-language models.

Proposing Cultural Awareness Score for image captions.

Creating dataset for culturally sensitive AI applications.

Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-language model evaluation

Cultural Awareness Score metric

MOSAIC-1.5k dataset introduction

🔎 Similar Papers

See It from My Perspective: How Language Affects Cultural Bias in Image Understanding