Safe at the Margins: A General Approach to Safety Alignment in Low-Resource English Languages -- A Singlish Case Study

📅 2025-02-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Mainstream large language models (LLMs) exhibit safety misalignment on low-resource English varieties—e.g., Singlish—due to overreliance on Western-centric English data. Method: We propose the first safety alignment framework tailored to non-standard English, introducing KTO-S, a stabilized variant of Kahneman–Tversky Optimization (KTO), the first application of KTO to safety fine-tuning for low-resource languages. We theoretically show that DPO implicitly enforces only weak safety objectives and empirically demonstrate that SFT+KTO significantly outperforms DPO. Results: Applied to SEA-Lion-v2.1-Instruct (a Llama3-8B variant), our method reduces toxicity by 99% on the Singlish safety benchmark, maintains strong generalization to TOXIGEN, and preserves full performance on standard academic benchmarks (MMLU, BBH), with no accuracy degradation.

Technology Category

Application Category

📝 Abstract
To ensure safe usage, Large Language Models (LLMs) typically undergo alignment with human-defined values. However, this alignment often relies on primarily English data and is biased towards Western-centric values, limiting its effectiveness in low-resource language settings. In this paper, we describe our approach for aligning SEA-Lion-v2.1-Instruct (a Llama3-8B variant) to minimize toxicity in Singlish, an English creole specific to Singapore. We find that supervised fine-tuning and Kahneman-Tversky Optimization (KTO) on paired and unpaired preferences is more sample efficient and yields significantly better results than Direct Preference Optimization (DPO). Our analysis reveals that DPO implicitly enforces a weaker safety objective than KTO, and that SFT complements KTO by improving training stability. Finally, we introduce a simple but novel modification to KTO, KTO-S, which improves training stability through better gradient exploitation. Overall, we present a general approach for safety alignment conducive to low-resource English languages, successfully reducing toxicity by 99% on our Singlish benchmark, with gains generalizing to the broader TOXIGEN dataset while maintaining strong performance across standard LLM benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Safety alignment in low-resource languages
Minimizing toxicity in Singlish
Improving training stability with KTO-S
Innovation

Methods, ideas, or system contributions that make the work stand out.

Supervised fine-tuning enhances training stability
Kahneman-Tversky Optimization improves safety alignment
KTO-S modification optimizes gradient exploitation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
I
Isaac Lim
GovTech Singapore
S
Shaun Khoo
GovTech Singapore
W
Watson Chua
GovTech Singapore
G
Goh Jiayi
GovTech Singapore
J
Jessica Foo
GovTech Singapore