Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入BLUEPRINT框架和WORLDVIEWSIM模块,利用蒙特卡洛树搜索优化对话策略,揭示了多轮对话中模型对特定影响因素的脆弱性及恢复路径。
📝 Abstract
Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module. Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory. Across six frontier models, BLUEPRINT achieves near-ceiling ASR on major open-weight and proprietary models, while requiring the fewest average queries (2.46). The resulting trajectories further reveal model-specific vulnerability among resistant targets: each responds to distinct influence factors and strategy transitions, yet all share a common recovery pathway-shifting toward concrete, executable task framing consistently escapes hard-refusal states. Ablations confirm operational cues matter most: making requests actionable has the largest impact, gain framing is unusually potent, and some legitimacy appeals can backfire. These findings suggest robust multi-turn safety requires monitoring not only harmful content, but also how dialogue state makes unsafe requests appear concrete and locally executable.
Problem

Research questions and friction points this paper is trying to address.

multi-turn jailbreak attacks
conversational mechanisms
vulnerability
Innovation

Methods, ideas, or system contributions that make the work stand out.

BLUEPRINT
WORLDVIEWSIM
Monte Carlo Tree Search
multi-turn dialogue
safety evaluation
S
Siyu Chen
AI4S Center, Shanghai Qi Zhi Institute, Shanghai, 200232, China
H
Haoran Wang
AI4S Center, Shanghai Qi Zhi Institute, Shanghai, 200232, China
X
Xiaojian Li
College of AI, Tsinghua University, Beijing, 100083, China; Fangcun AI, Beijing, 100084, China
Yao Huang
Yao Huang
Institute of Artificial Intelligence, Beihang University
Trustworthy MLMultimodal Learning
Yinpeng Dong
Yinpeng Dong
Tsinghua University
Machine LearningDeep LearningAI Safety
W
Wei Xu
AI4S Center, Shanghai Qi Zhi Institute, Shanghai, 200232, China; College of AI, Tsinghua University, Beijing, 100083, China; Institute for Interdisciplinary Information Sciences, Tsinghua University, Beijing, 100084, China