Is Multilingual LLM Watermarking Truly Multilingual? A Simple Back-Translation Solution

📅 2025-10-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing multilingual watermarking methods for large language models exhibit fragility in low- and medium-resource languages due to semantic clustering failure and are vulnerable to translation attacks, undermining cross-lingual traceability. This paper proposes STEAM, a non-intrusive watermark enhancement method based on back-translation that restores watermark signals diluted by translation via bilingual alignment and semantic consistency verification. STEAM requires no model training, is compatible with any existing watermarking scheme, and supports multiple tokenizers and rapid extension to new languages. Experiments across 17 languages demonstrate that STEAM improves average AUC by 0.19 and TPR@1% by 40 percentage points, significantly enhancing watermark robustness and traceability in low- and medium-resource languages. To our knowledge, STEAM is the first approach to achieve truly cross-lingual, highly robust multilingual content provenance.

Technology Category

Application Category

📝 Abstract
Multilingual watermarking aims to make large language model (LLM) outputs traceable across languages, yet current methods still fall short. Despite claims of cross-lingual robustness, they are evaluated only on high-resource languages. We show that existing multilingual watermarking methods are not truly multilingual: they fail to remain robust under translation attacks in medium- and low-resource languages. We trace this failure to semantic clustering, which fails when the tokenizer vocabulary contains too few full-word tokens for a given language. To address this, we introduce STEAM, a back-translation-based detection method that restores watermark strength lost through translation. STEAM is compatible with any watermarking method, robust across different tokenizers and languages, non-invasive, and easily extendable to new languages. With average gains of +0.19 AUC and +40%p TPR@1% on 17 languages, STEAM provides a simple and robust path toward fairer watermarking across diverse languages.
Problem

Research questions and friction points this paper is trying to address.

Existing multilingual watermarking fails in medium- and low-resource languages
Current methods lack robustness under translation attacks across languages
Tokenizer vocabulary limitations cause semantic clustering failures in some languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Back-translation detection method restores watermark strength
STEAM is compatible with any watermarking technique
Robust across different tokenizers and languages
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
A
Asim Mohamed
African Institute for Mathematical Sciences
M
Martin Gubri
Parameter Lab