Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis

๐Ÿ“… 2026-06-08
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the limitations of current safety evaluations for multilingual large language models, which often rely on literal translations of English benchmarks and overlook cultural contextual differences, leading to inaccurate risk assessments. The authors present the first systematic construction of paired red-teaming datasets in Korean, Japanese, Thai, and Khmer, featuring both direct translation (DT) and culturally adapted (CA) prompts matched one-to-one by seed. Evaluation using attack success rate (ASR) and a novel Cultural Contextualization Score (C3) reveals that CA prompts increase ASR by an average of 9.3 percentage points across all 16 languageโ€“model combinations. Furthermore, DT significantly underestimates local threats in 44 out of 48 risk categories, while C3 scores rise from a mean of 0.17 to as high as 2.51, demonstrating that cultural adaptation is essential for accurately capturing localized safety risks.
๐Ÿ“ Abstract
Multilingual safety evaluation of large language models (LLMs) has predominantly relied on direct translation (DT) of English benchmarks into target languages - an approach that converts surface-level linguistic form while failing to reflect the cultural context embedded in threat scenarios, social norms, and legal frameworks. We construct paired DT and culturally-adapted (CA) datasets via 1:1 seed matching for four languages - Korean (KO), Japanese (JA), Thai (TH), and Khmer (KM) - and compare Attack Success Rate (ASR) and Cultural Realism scores across four open-source LLM. CA prompts yield Delta-ASR > 0 across all 16 language x model combinations (mean +9.3 pp), and DT-based evaluation underestimates risk in 44 of 48 category x language combinations. Language-level analysis reveals that the distribution of threat forms is heterogeneous across languages. Cultural Realism analysis further shows that DT Cultural Depth (C3) scores remain consistently below 1.0 out of 3.0 across all four languages (mean 0.17), whereas CA scores reach up to 2.51, indicating that direct translation produces inputs systematically divergent from those encountered in real-world multicultural settings. These findings demonstrate that adapting benchmarks to language-specific cultural contexts - rather than relying on linguistic translation alone - is necessary for valid multilingual LLM safety evaluation.
Problem

Research questions and friction points this paper is trying to address.

multilingual safety evaluation
large language models
cultural adaptation
direct translation
red-teaming
Innovation

Methods, ideas, or system contributions that make the work stand out.

culturally-adapted red-teaming
multilingual LLM safety
direct translation limitation
cultural realism
Attack Success Rate (ASR)
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
H
Hyeji Choi
AI Safety Team, DATUMO.INC, Seoul, South Korea
Y
Yongtaek Lim
AI Safety Team, DATUMO.INC, Seoul, South Korea
M
Minwoo Kim
AI Safety Team, DATUMO.INC, Seoul, South Korea