Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety Moderator

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决AI应用中内容安全审核问题,提出Nemotron 3.5 CS模型,该模型能跨12种语言处理文本、图像和文档,并提供自定义策略支持。
📝 Abstract
Safety moderation for deployed AI applications is moving beyond text-only prompts: systems increasingly need to judge images, documents, screenshots, and generated responses under policies that vary across domains. Existing guardrails usually cover only part of this setting, making it difficult to combine broad coverage, custom policy control, and low compute cost. We present Nemotron 3.5 Content Safety Moderator, also referred to as Nemotron 3.5 CS in this paper for brevity, a compact 4B vision-language safety moderator that jointly classifies user prompts, images, and assistant responses across 12 languages. Nemotron 3.5 CS returns safety labels for latency-sensitive moderation and can additionally produce concise reasoning traces that apply supplied custom policies and identify violated categories when reasoning is requested. We also release a multimodal and multilingual safety dataset for guard training, spanning human-labeled real-image moderation, benign vision-language and document tasks, synthetic rare-risk and jailbreak cases, and custom-policy examples. Across evaluations spanning multimodal safety, text moderation, multilingual robustness, custom-policy following, benign false positives, and latency, Nemotron 3.5 CS demonstrates a practical coverage tradeoff: it adds image-conditioned and policy-conditioned moderation while remaining broadly competitive with specialized guard models. These results suggest that compact vision-language moderators can serve as deployable front-line safety components, with reasoning used selectively for audit and policy review.
Problem

Research questions and friction points this paper is trying to address.

multimodal
multilingual
content safety
policy control
low compute cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal
Multilingual
Reasoning Enabled
Compact Model
Custom Policy Control
🔎 Similar Papers
No similar papers found.
V
Varun Singh
NVIDIA, Santa Clara, CA
A
Anuj Doshi
NVIDIA, Santa Clara, CA
Makesh Narsimhan Sreedhar
Makesh Narsimhan Sreedhar
NVIDIA
Language ModelsDialog AgentsMachine TranslationNatural Language Processing
S
Shaona Ghosh
NVIDIA, Santa Clara, CA
K
Katherine Luna
NVIDIA, Santa Clara, CA