Playing Whack-a-Mole with misconceptions about memorization, extraction, and copyright

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该文指出通过微调记忆化结果来解决版权问题的方法存在测量程序无效、缺乏对照实验等问题,其声称的从训练数据中提取大量版权书籍内容的说法未得到支持。
📝 Abstract
After careful review, I'm confident the headline fine-tuning memorization results in Alignment Whack-a-Mole use an invalid measurement procedure. The book memorization coverage metric these headline results depend on counts sequence matches far shorter than what field standards consider valid evidence of memorization, and the prompting procedure used to elicit memorization runs the risk of leaking the text being "extracted" in the prompt. The paper doesn't include the negative-control experiments needed to see how much the results are inflated by false positives: claiming extraction success (and therefore memorization of training data) when matches between generations and training data may be due to other factors. Given these validity issues, the paper's claims that fine-tuning lets users extract substantial portions of copyrighted books, in a form that could substitute for the originals, aren't supported by the reported results. The failure to report the experiments' cost (an important component of the threat model) further compromises the copyright claims. I'm writing this note because, in the last month, (prospective) plaintiffs have reached out to me to ask about this paper. They're looking to cite this work as valid evidence in support of claims in ongoing and potential future copyright litigation.
Problem

Research questions and friction points this paper is trying to address.

memorization
extraction
copyright
fine-tuning
measurement procedure
Innovation

Methods, ideas, or system contributions that make the work stand out.

memorization
false positives
copyright claims
negative-control experiments
measurement procedure
💼 Related Jobs
No related jobs found.