Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出一种新的白盒攻击方法,通过结合知识编辑和关联上下文检索技术,诱导大型语言模型产生不安全行为,实验表明该方法有效且不影响模型整体性能。
📝 Abstract
As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. We propose a novel white-box attack inspired by locate-then-edit approaches from the field of Knowledge Editing. Our choice is motivated by the observation that models edited with such schemes tend to assign unusually high prediction probabilities to the edit target, a property that is particularly advantageous when designing attacks. We modify the editing framework by incorporating as- sociative knowledge retrieved from the model, thereby extending constraint removal to an entire thematic category rather than being limited to prompts from a predefined dataset. Experiments with various archi- tectures demonstrate improved attack effectiveness over competing methods without dealing critical damage to general model performance.
Problem

Research questions and friction points this paper is trying to address.

large language models
unsafe behavior
white-box attack
Innovation

Methods, ideas, or system contributions that make the work stand out.

Association Context Retrieval
Knowledge Editing
White-Box Attacks
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Roman Maksimov
Basic Research of Artificial Intelligence Laboratory
V
Vladimir Aletov
Basic Research of Artificial Intelligence Laboratory
Vladimir Solodkin
Vladimir Solodkin
MBZUAI
OptimizationReinforcement Learning
D
Dmitry Bylinkin
Basic Research of Artificial Intelligence Laboratory
Daniil Medyakov
Daniil Medyakov
Unknown affiliation
Optimization
Aleksandr Beznosikov
Aleksandr Beznosikov
PhD, Basic Research of Artificial Intelligence Lab
OptimizationMachine Learning