EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出EvoFlint,通过进化多样性搜索方法构建多轮对话攻击策略档案,以探索和细化大型语言模型的漏洞。
📝 Abstract
Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-teaming methods treat this as a generation problem: produce attacks that break the model. We argue it is better framed as a search problem: discover, organize, and iteratively refine a diverse archive of attack strategies, producing a structured map of how a target model fails rather than a list of one-off successes. We introduce EvoFlint, which applies evolutionary quality-diversity search to multi-turn red-teaming. Attack strategies are phased conversation plans, not raw prompts, and are evolved through LLM-driven mutation and crossover. A Pareto fitness over attack success rate and peak severity preserves selection signal from near-miss attacks. A risk-indexed archive runs novelty search with local competition over strategy description embeddings inside each cell, maintaining diversity without committing to a predefined style taxonomy. A generation-level memory accumulates target-model insights across the population and feeds them back into strategy generation. On the HarmBench-test split, EvoFlint reaches attack success rates of 35.8% on Claude Sonnet 4.6, 59.7% on GPT-5.4, and 94.3% on Qwen3-32B, alongside 98.7% on the older GPT-4o included as a baseline reference. The resulting archive, organized by risk category, exposes for each target which categories of harm its safety training has and has not covered.
Problem

Research questions and friction points this paper is trying to address.

multi-turn attacks
language models
vulnerabilities
harmful prompts
Innovation

Methods, ideas, or system contributions that make the work stand out.

evolutionary quality-diversity search
multi-turn red-teaming
Pareto fitness
novelty search with local competition
generation-level memory
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
F
Feitong Qiao
Reinforce Labs, USA
L
Liren Peng
Reinforce Labs, USA
S
Shiming Ren
Reinforce Labs, USA
Aishwarya Jadhav
Aishwarya Jadhav
Autopilot Engineer at Tesla; Masters Student at Carnegie Mellon University
A
Arghavan Bahadorinejad
Reinforce Labs, USA
M
Marinette Chen
Reinforce Labs, USA
Muhan Zhang
Muhan Zhang
Peking University
Machine LearningGraph Neural NetworkLarge Language Models
A
Abdulaziz Suria
Reinforce Labs, USA
G
Gennevi Lu
Reinforce Labs, USA
A
Anish Das Sarma
Reinforce Labs, USA