Stress Testing Unlearning Algorithms

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对大语言模型中难以彻底删除特定训练数据影响的问题,提出了一种新的评估基准WMDP++,通过主动测试信息提取和边界问题性能来改进现有方法。
📝 Abstract
Recently, machine unlearning, the removal of specific training data influence from a model, has gained increasing attention. In large language models (LLMs), unlearning is particularly challenging due to the ambiguity of inputs and outputs. Con- sequently, rigorous evaluation is critical for assessing both safety and utility, and for driving progress in unlearning meth- ods. We identify two key shortcomings in existing unlearning benchmarks: (1) they do not actively test whether unlearned information can still be forcibly extracted, and (2) they fail to evaluate performance preservation on boundary questions, be- nign queries that are semantically close to the unlearned con- tent. Here we introduce WMDP++, an extension of WMDP that addresses these gaps by incorporating targeted extrac- tion of unlearned information and systematic evaluation on boundary questions. WMDP++ provides a more stringent and informative benchmark for evaluating unlearning in LLMs.
Problem

Research questions and friction points this paper is trying to address.

machine unlearning
large language models
evaluation benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

machine unlearning
targeted extraction
boundary questions
performance preservation
N
Noam Diamant
Bar-Ilan University
Ethan Fetaya
Ethan Fetaya
Bar-Ilan University
Machine learningComputer vision
N
Neta Glazer
Bar-Ilan University