The Price of Consistency: Exploiting Visual Anchors for Multimodal Jailbreaking in Video Generation

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究揭示了视觉锚定效应增加视频生成安全性风险,提出DIVA框架利用此漏洞,通过分离有害意图和动态文本提示实现攻击。
📝 Abstract
The rapid evolution of video generation has shifted the paradigm from pure text-driven to multi-conditional controllable generation, with reference images now widely adopted as conditional inputs to achieve superior spatiotemporal consistency. While these reference images serve as powerful visual anchors that significantly enhance controllability, their impact on safety remains largely unexplored. In this work, we reveal the visual anchoring effect: by enforcing consistency, the mechanism prevents the generated content from drifting away from the original harmful intent, thereby eliminating the model's natural safety escape route from harmful to benign content. Consequently, visual anchors inherently increase the safety risk---this is the price of consistency. Building on this insight, we propose Decoupling Intent via Visual Anchors (DIVA), a training-free multimodal jailbreak framework for video generation that exploits this vulnerability. DIVA decouples harmful intent into a static visual anchor image and a dynamic motion text prompt, and employs dual-criteria selection to balance attack stealthiness with semantic preservation. Extensive experiments across various leading commercial platforms and mainstream open-source video generation models demonstrate that DIVA achieves a substantially higher Attack Success Rate than existing text-only methods. To facilitate future research, we additionally contribute TI2VSafetyBench, the first safety benchmark for multi-conditional video generation.
Problem

Research questions and friction points this paper is trying to address.

visual anchors
safety risk
consistency
video generation
harmful intent
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual anchoring effect
multimodal jailbreak
DIVA
dual-criteria selection
🔎 Similar Papers
P
Peng Li
School of Computer Science and Technology, University of Chinese Academy of Sciences
Q
Qianqian Xu
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences
Yangbangyan Jiang
Yangbangyan Jiang
University of Chinese Academy of Sciences
Machine LearningDeep learning
Z
Zhipeng Yu
School of Computer Science (National Pilot Software Engineering School), Beijing University of Posts and Telecommunications
Qingming Huang
Qingming Huang
University of the Chinese Academy of Sciences
Multimedia Analysis and RetrievalImage and Video ProcessingPattern RecognitionComputer VisionVideo Coding