Arabic Safety Alignment as Selective Refusal: An Empirical Study of SFT, DPO, and Guard Calibration

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过监督微调(SFT)、直接偏好优化(DPO)和推理防护方法,解决阿拉伯语大模型在拒绝有害提示同时不过度拒绝良性或敏感提示的问题。
📝 Abstract
Arabic large language models must refuse harmful prompts without over-refusing benign or sensitive prompts, yet a single refusal rate hides this trade-off. We evaluate it using benign refusal B and harmful-prompt refusal H, where H measures refusal rather than harmful compliance. Across five Arabic-capable models and 130 runs on the full human-written AraSafe set, refusal-only supervised fine-tuning (SFT) collapses toward blanket refusal, whereas selected mixed-SFT configurations reach H = 90% to 93% at B = 14% to 23%; four selected configurations exceed the H = 90% target in all three runs, while Fanar does so in two of three. Direct Preference Optimization (DPO) and inference guards change B and H differently across models rather than acting as uniform upgrades. In a blinded 300-response audit, annotator binary-refusal agreement is 89.0% (kappa = 0.78); Qwen3Guard and Aya Expanse 32B reach 88.7% and 91.0% accuracy, respectively, with no conclusive paired difference. Selected SFT raises H on Arabizi for all five models, but none reaches 90%, showing only partial transfer from Modern Standard Arabic. Overall, the results support model-specific operating-point selection: set a deployment target and retain only interventions that improve it.
Problem

Research questions and friction points this paper is trying to address.

Arabic large language models
harmful prompts
benign prompts
refusal rate
trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective Refusal
Supervised Fine-Tuning (SFT)
Direct Preference Optimization (DPO)
Guard Calibration
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Mohamad Zbib
American University of Beirut
A
Ammar Mohanna
American University of Beirut