ForgetMark: Stealthy Fingerprint Embedding via Targeted Unlearning in Language Models
This work proposes a covert fingerprinting mechanism based on directed forgetting to address the vulnerability of existing language model watermarking methods to filtering, heuristic detection, and false triggers. By leveraging an auxiliary model and prediction entropy ranking, the method constructs compact key-value pairs and trains a lightweight LoRA adapter to selectively suppress original responses associated with specific keys, embedding imperceptible forgetting traces without compromising the model’s general capabilities. Departing from conventional fixed trigger-response paradigms, it exploits probabilistic forgetting patterns to significantly enhance stealth and reduce false positives. Combining likelihood and semantic evidence, the approach achieves 100% ownership verification accuracy under black-box and gray-box settings, remains robust against model merging and incremental fine-tuning, incurs no performance degradation on standard tasks, and consistently outperforms backdoor-based baselines.