Evaluating Perspectival Biases in Cross-Modal Retrieval

๐Ÿ“… 2025-10-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work identifies two novel types of perspective bias in cross-modal retrieval: *linguistic popularity bias*โ€”where image-to-text retrieval favors high-frequency linguistic items over semantically optimal onesโ€”and *cultural associativity bias*โ€”where text-to-image retrieval prefers culturally stereotypical images over semantically accurate ones. We propose a multilingual and multicultural empirical framework grounded in cross-modal alignment analysis, systematically evaluating bias origins and mitigation strategies across a dataset spanning 12 languages and six cultural regions. Our findings show that explicit alignment significantly mitigates popularity bias but yields limited improvement for cultural associativity bias, indicating its entrenchment in deep semantic representations and necessitating structured interventions beyond data augmentation. This study is the first to formally define, distinguish, and comparatively analyze these two biases, establishing cultural associativity bias as more persistent and challenging. The work provides both theoretical foundations and methodological paradigms for developing fairer cross-modal retrieval systems.

Technology Category

Application Category

๐Ÿ“ Abstract
Multimodal retrieval systems are expected to operate in a semantic space, agnostic to the language or cultural origin of the query. In practice, however, retrieval outcomes systematically reflect perspectival biases: deviations shaped by linguistic prevalence and cultural associations. We study two such biases. First, prevalence bias refers to the tendency to favor entries from prevalent languages over semantically faithful entries in image-to-text retrieval. Second, association bias refers to the tendency to favor images culturally associated with the query over semantically correct ones in text-to-image retrieval. Results show that explicit alignment is a more effective strategy for mitigating prevalence bias. However, association bias remains a distinct and more challenging problem. These findings suggest that achieving truly equitable multimodal systems requires targeted strategies beyond simple data scaling and that bias arising from cultural association may be treated as a more challenging problem than one arising from linguistic prevalence.
Problem

Research questions and friction points this paper is trying to address.

Evaluating perspectival biases in cross-modal retrieval systems
Analyzing prevalence bias favoring dominant languages in retrieval
Examining association bias prioritizing culturally linked images
Innovation

Methods, ideas, or system contributions that make the work stand out.

Explicit alignment mitigates prevalence bias effectively
Association bias remains a distinct challenging problem
Targeted strategies needed beyond simple data scaling
๐Ÿ”Ž Similar Papers
No similar papers found.
T
Teerapol Saengsukhiran
Department of Computer Engineering, Chulalongkorn University, Bangkok, Thailand 10330
P
Peerawat Chomphooyod
Department of Computer Engineering, Chulalongkorn University, Bangkok, Thailand 10330
N
Narabodee Rodjananant
Department of Computer Engineering, Chulalongkorn University, Bangkok, Thailand 10330
C
Chompakorn Chaksangchaichot
Department of Computer Engineering, Chulalongkorn University, Bangkok, Thailand 10330
P
Patawee Prakrankamanant
Department of Computer Engineering, Chulalongkorn University, Bangkok, Thailand 10330
W
Witthawin Sripheanpol
Department of Computer Engineering, Chulalongkorn University, Bangkok, Thailand 10330
P
Pak Lovichit
Department of Computer Engineering, Chulalongkorn University, Bangkok, Thailand 10330
Sarana Nutanong
Sarana Nutanong
Vidyasirimedhi Institute of Science and Technology
Natural Language Processing and Representation Learning
Ekapol Chuangsuwanich
Ekapol Chuangsuwanich
Chulalongkorn University
Speech ProcessingNatural Language ProcessingMedical AI