SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过构建大规模五级安全评估数据集SafeAtlas-VL及相应的守卫模型,解决了多模态交互中难以比较和识别风险的问题。
📝 Abstract
Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for a single judgment target and reduce safety assessment to a binary decision. Consequently, risk becomes difficult to compare across a multimodal interaction, and ambiguous cases are obscured. We introduce SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale. We curate a broad collection of safety-relevant data from both real-world and synthetic sources and apply a disagreement-aware annotation procedure. The resulting dataset spans 15 harm categories and 55 fine-grained subcategories, covering a broad range of multimodal safety scenarios. We also construct SafeAtlas-Bench, a held-out set of 5,000 instances for evaluating five-level predictions and continuous risk scores. Upon this dataset, we train the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection. Our models not only perform five-way classification of safety levels but also map safety to continuous scores through a soft cumulative ordinal head. Experimental results demonstrate that guard models trained on our dataset exhibit strong generalization: even without using the training sets of other benchmarks, they achieve competitive performance on the corresponding test sets. Notably, our 8B model attains the overall best performance, outperforming the previous SOTA by approximately 4% in F1 score. Code, data, and models are released to support further research. Warning: this paper contains example data that may be offensive, harmful, graphic, or disturbing.
Problem

Research questions and friction points this paper is trying to address.

multimodal safety
risk assessment
binary decision
ambiguous cases
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal safety
five-level scale
target-conditioned tuning
continuous risk scores
generalization
Z
Zongrui Wang
Shanghai Jiao Tong University
X
Xiangyang Zhu
Shanghai Artificial Intelligence Laboratory
S
Sicheng Wang
Shanghai Artificial Intelligence Laboratory
H
Han Wang
Shanghai Artificial Intelligence Laboratory
D
Dingyi Rong
Shanghai Jiao Tong University
Z
Zeyu Zhang
Shanghai Jiao Tong University
Chunyi Li
Chunyi Li
NTU | SJTU | Shanghai AI Lab
Generative AIEmbodied AILow-level Vision
Y
Yue Shi
Shanghai Artificial Intelligence Laboratory
K
Kaiwei Zhang
Shanghai Artificial Intelligence Laboratory
Z
Zicheng Zhang
Shanghai Artificial Intelligence Laboratory
Y
Yuan Tian
Shanghai Artificial Intelligence Laboratory
Q
Qi Jia
Shanghai Artificial Intelligence Laboratory
Y
Yan Teng
Shanghai Artificial Intelligence Laboratory
W
Wei Sun
East China Normal University
N
Ning Liu
Shanghai Jiao Tong University
Guangtao Zhai
Guangtao Zhai
Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI EvaluationDisplays