Selective Knowledge Edit Reversal via Gated Singular Vector Shrinkage

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对大语言模型中恶意编辑带来的安全风险问题,提出了一种基于谱的方法来选择性地逆转特定编辑,同时保留其他有益的编辑。
📝 Abstract
Knowledge editing provides an efficient way to update factual knowledge in large language models. However, malicious edits may introduce safety risks, making it necessary to reverse undesirable editing effects. Existing reversal methods for parameter-modifying edits mainly focus on global removal, which may also erase beneficial edits that should be preserved. In this paper, we study selective reversal of edited knowledge, where the goal is to reverse targeted edited facts while preserving the remaining edited facts. Based on the hypothesis that each edit is sparsely encoded within the dominant subspace of the edited matrix, we propose a spectral-based reversal framework that locates edit-sensitive components within the dominant singular subspace of edited weights. Experiments across multiple settings demonstrate the effectiveness of our method in reversing selected edits while preserving unrelated edited facts. These results suggest that different edits are sparsely encoded within dominant singular components and can be separable when the number of edits is moderate, making selective spectral reversal a promising direction for locating edit-specific components and repairing edited language models.
Problem

Research questions and friction points this paper is trying to address.

knowledge editing
safety risks
selective reversal
edited facts
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective Reversal
Spectral-based Framework
Gated Singular Vector Shrinkage
Edit-sensitive Components
💼 Related Jobs
No related jobs found.
W
Weifeng Jiang
College of Computing and Data Science, Nanyang Technological University, Singapore.
R
Ruirui Chen
Institute of Advanced Intelligence and Computing, Agency for Science, Technology and Research, Singapore.
Qianren Mao
Qianren Mao
Zhongguancun Laboratory
Text miningText GenerationKnowledge Graph and Reasoing
J
Junnan Liu
Department of Data Science and AI, Faculty of Information Technology, Monash University, Australia.
Q
Qili Zhang
Zhongguancun Laboratory, China.
Kwok-Yan Lam
Kwok-Yan Lam
Nanyang Technological University
CybersecurityPrivacy-Preserving technologiesDigital TrustDistributing systemsLegalTech