LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
Emerging safety risks in multi-turn multimodal (MMT) dialogue with vision-language models (VLMs) involve malicious intent distributed stealthily across turns and modalities, evading detection by single-turn or single-modality safety auditing. Method: This work formally defines the MMT dialogue safety problem; introduces MMDS—the first fine-grained, evidence-annotated, safety-specific dataset for MMT dialogue; proposes a Monte Carlo Tree Search–based multimodal red-teaming framework to automatically generate cross-turn, cross-modal harmful dialogues; and designs a risk assessment model integrating multimodal contextual modeling and policy-aware reasoning for joint input-response safety evaluation. Contribution/Results: The proposed LLaVAShield framework achieves significant improvements over strong baselines on multi-turn content moderation tasks and maintains state-of-the-art performance under dynamic safety policies.