Auditing Cross-Lingual Fairness in Language Model Watermarking

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种评估框架,通过校准检测阈值、独立测量及三种质量度量方法,解决跨语言环境下语言模型水印公平性审计问题。
📝 Abstract
Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually. We propose an evaluation framework with four components: detection thresholds calibrated empirically per deployment context, a threshold-independent companion measurement that distinguishes calibration failures from detection failures, three disjoint quality measurement paradigms (distributional, paired-semantic, and reference-perplexity), and a generalized-entropy decomposition of cross-language disparity over a typological family partition. Applied to six watermarking schemes, three open-weight generators, eleven languages spanning four scripts and eight typological families, and both base and instruction-tuned regimes, the framework reveals failure modes that single-language single-paradigm evaluation cannot surface. Across detection and quality, observed disparity is predominantly between-family on the typological partition, indicating that cross-lingual fairness gaps in watermarking are structural to language properties rather than idiosyncratic to particular languages.
Problem

Research questions and friction points this paper is trying to address.

Cross-Lingual Fairness
Language Model Watermarking
Multilingual Deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-lingual fairness
evaluation framework
watermarking schemes
typological family