Optimal Watermark Localization in Mixed-Source Large Language Model Texts

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of localizing text watermarks in large language model outputs with mixed sources by proposing an adaptive thresholding method based on pivot statistics. By establishing a theoretical phase transition boundary for watermark localization and integrating multiple testing corrections with a prior-free adaptive estimation mechanism, the approach effectively mitigates signal interference. Theoretically, the method achieves the optimal detection boundary and demonstrates precise localization even under editing attacks. These contributions significantly enhance the robustness and accuracy of watermark detection in complex scenarios, providing reliable theoretical and technical support for copyright protection of hybrid-source texts generated by large language models.
📝 Abstract
Watermarking provides a principled way to authenticate text generated by large language models (LLMs). In practice, however, the final text may be mixed-source, with watermark evidence surviving at only a subset of token positions after rewriting, insertion, deletion, or paraphrasing. Although prior work has studied global detection of watermark signals, when such signals can be localized remains unclear. We formulate watermark localization as a token-level multiple-testing problem based on pivotal statistics, with a latent indicator recording whether watermark dependence survives at each position. Under an asymptotic regime indexed by exponents for signal sparsity, next-token concentration, and effective-vocabulary growth, we derive a sharp boundary for global detection and phase transitions for discovery and classification within the class of coordinatewise pivot-based localization rules. We show that discovery is strictly harder than detection and that consistent classification is impossible across the parameter regime within this class. We then develop an adaptive thresholding method that does not require knowledge of the exponents or time-varying next-token distributions, but uses a data-driven estimate of the surviving watermark fraction. The method attains the optimal discovery boundary and near-optimal discovery power relative to homogeneous pivot-based rules. Simulations support the theoretical phase transitions, while experiments on model-generated texts demonstrate practical localization performance under common edit mechanisms.
Problem

Research questions and friction points this paper is trying to address.

Watermark Localization
Large Language Models
Mixed-Source Text
Token-level Detection
Multiple Testing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Watermark Localization
Multiple Testing
Adaptive Thresholding
Phase Transitions
Mixed-Source Text