Redteaming Leading Arabic LLMs with ASAS

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决阿拉伯语大语言模型的安全性和文化适应性问题,通过创建ASAS基准进行红队评估,揭示了模型在高风险类别中的安全漏洞。
📝 Abstract
As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasingly critical. However, Arabic LLM safety remains underexplored, especially in adversarial evaluation settings. We introduce the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming LLMs. ASAS contains 801 prompts spanning 8 safety categories and 8 attack strategies, with ideal responses in Modern Standard Arabic (MSA). We conduct a redteaming evaluation across seven leading models with Arabic capabilities, including GPT-4o, Claude 3.7 Sonnet, and regional models such as ALLaM and FANAR. Human annotators rate responses using a structured 4-point safety scale, revealing that most models fail to defend against 50% of unsafe prompts. Our findings highlight major safety gaps in high-harm categories such as weapons and illicit substances, with direct and obfuscation-based attacks proving most effective. The results also show that language alignment does not readily transfer across languages, and that automated safety judges (e.g., GPT-4o) perform poorly compared to human annotators. ASAS provides a culturally grounded benchmark and redteaming protocol to drive progress in Arabic LLM safety.
Problem

Research questions and friction points this paper is trying to address.

Arabic LLMs
safety
cultural alignment
adversarial evaluation
redteaming
Innovation

Methods, ideas, or system contributions that make the work stand out.

Arabic Safety Index
Redteaming
LLM Safety
Cultural Alignment
🔎 Similar Papers
No similar papers found.