MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对多语言长文本LLM生成归因问题,提出MultiGhostBench基准,包含多种语言和脚本的928本书,用于评估不同方法在分布偏移下的表现。
📝 Abstract
While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relatively short texts. We introduce MultiGhostBench, a multilingual benchmark comprising 928 books generated by five recent LLMs across six languages and three scripts, with an average length of approximately 59K words per book. The benchmark supports evaluation under domain, author, and language shifts. Evaluation of representative AA methods shows that no single method consistently performs best across settings, and performance generally degrades under distribution shifts. Transformer-based detectors can retain generator-related information across languages, although transfer effectiveness varies by language pair, whereas statistical and fingerprint-based detectors are more language-dependent. We envision MultiGhostBench as a valuable resource for the development and evaluation of robust AA methods. The dataset and code can be found at https://github.com/GrecoMT/MultiGhostBench.
Problem

Research questions and friction points this paper is trying to address.

multilingual
authorship attribution
distribution shifts
long-form text
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multilingual Benchmark
Long-Form LLM-Generated Text
Distribution Shifts
Authorship Attribution