The Illusion of Cross-Lingual Safety in Low-Resource Languages
Current safety alignment of large language models is predominantly based on English, and their cross-lingual generalization to low-resource languages remains poorly understood, posing potential risks. This work introduces the LoDNA dataset, comprising both literal translations and culturally localized prompts, to systematically evaluate the transferability of safety mechanisms across four African languages. We propose a probing method grounded in the geometric structure of the model’s latent space to analyze internal representations underlying refusal behaviors. Our study reveals, for the first time, significant limitations in cross-lingual safety alignment: in most language–model combinations, harmful prompts retain less than 10% of the refusal signal observed in English, indicating that semantic alignment does not ensure consistent safety routing. These findings challenge the assumption of a language-invariant harm manifold.