Chainwash: Multi-Step Rewriting Attacks on Diffusion Language Model Watermarks
This study addresses the vulnerability of existing language model watermarks under iterative rewriting attacks, which significantly degrade detection reliability. The work systematically evaluates the robustness of watermarking schemes in diffusion-based language models subjected to multi-round, multi-style chained rewrites—including paraphrasing, simplification, and academic stylization—and reveals, for the first time, the cumulative destructive effect of sequential rewrites on watermark integrity. Experiments employ four open-source large language models (1.5B–8B parameters) within a statistical watermark detection framework, subjecting generated text to five rounds of chained rewrites. Results demonstrate a sharp decline in detection efficacy: from an initial rate of 87.9% to 14%–41% after a single rewrite, and further down to merely 4.86% after five rounds, establishing multi-step rewriting as a highly effective watermark removal attack.