The Illusion of Cross-Lingual Safety in Low-Resource Languages

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current safety alignment of large language models is predominantly based on English, and their cross-lingual generalization to low-resource languages remains poorly understood, posing potential risks. This work introduces the LoDNA dataset, comprising both literal translations and culturally localized prompts, to systematically evaluate the transferability of safety mechanisms across four African languages. We propose a probing method grounded in the geometric structure of the model’s latent space to analyze internal representations underlying refusal behaviors. Our study reveals, for the first time, significant limitations in cross-lingual safety alignment: in most language–model combinations, harmful prompts retain less than 10% of the refusal signal observed in English, indicating that semantic alignment does not ensure consistent safety routing. These findings challenge the assumption of a language-invariant harm manifold.
📝 Abstract
Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages. We investigate cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that pairs literal translations with culturally localized prompts. To move beyond generation-based evaluation, we propose a latent geometric framework that probes hidden-state refusal representations in LLMs. Our experimental results show that cross-lingual safety transfer is severely limited; harmful prompts retain less than 10% of the English refusal signal across most language-model pairs. Literal and localized prompts are semantically aligned (cosine 0.95-0.996) but drift across layers, suggesting models encode the concepts without routing them to safety mechanisms. These findings demonstrate that current multilingual safety alignment is superficial, providing strong evidence against the assumption of a universal, language-agnostic harm manifold within the specific low-resource languages studied. Warning: This paper contains example data that may be offensive or harmful.
Problem

Research questions and friction points this paper is trying to address.

cross-lingual safety
low-resource languages
safety alignment
multilingual LLMs
harm detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-lingual safety transfer
latent geometric framework
safety alignment
low-resource languages
refusal representations
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Abigail Oppong
Makerere University Center for Artificial Intelligence
P Sam Sahil
P Sam Sahil
Research Intern @ University of Hamburg
Machine LearningArtificial IntelligenceDeep LearningNLPComputer vision
Tadesse Destaw Belay
Tadesse Destaw Belay
Ph.D. candidate IPN, Mexico
NLP for Low-resource languagesMachine learningand LLMs
M
Maryam Ibrahim Mukhtar
Bayero University, Kano
E
Esmael Ahmed Abdu
Wollo University
Tassallah Abdullahi
Tassallah Abdullahi
Brown University
Natural Language ProcessingInformation RetrievalDigital Health
J
Jessica Oparebea
University of Ghana
S
Saminu Mohammad Aliyu
Bayero University, Kano
Idris Abdulmumin
Idris Abdulmumin
Postdoctoral Fellow, DSFSI, University of Pretoria
Machine TranslationNeural Machine TranslationNatural Language ProcessingInternet Technology
A
Abubakar Juma Chilala
Carnegie Mellon University
N
Nicholaus Dismas Ladislaus
Carnegie Mellon University
A
Alfred Malengo Kondoro
Hanyang University
L
Lemofouet Valdini Douglace
AIMS Cameroon
Shamsuddeen Hassan Muhammad
Shamsuddeen Hassan Muhammad
Bayero University, Kano, & Google DeepMind Academic Fellow at Imperial College London
Natural Language ProcessingSentiment AnalysisAfricaNLPLow-resource NLPMultilinguality
S
Seid Muhie Yimam
University of Hamburg