AuthBench: A Large-Scale Multilingual Benchmark for Authorship Representation across Genres and Lengths

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决现有作者身份基准的局限性,本文提出了AuthBench,一个大规模多语言基准,通过涵盖多种语言、体裁和文档长度来评估作者身份表示方法的有效性。
📝 Abstract
Authorship signals matter in settings where writing style carries identity: digital forensics, plagiarism analysis, account linking, misinformation investigation, and machine-generated text detection. Yet current authorship benchmarks remain fragmented, usually covering only a narrow language set, a single genre, or a limited document-length regime, which makes it difficult to assess whether modern representations truly generalize. We introduce AuthBench, a large-scale multilingual benchmark for authorship representation that is designed to make this evaluation broad, standardized, and realistic. AuthBench contains 428,150 documents written by 153,825 individuals across ten widely used languages, 9 primary genres, 66 fine-grained genres, and four document-length buckets. It supports two complementary tasks: authorship attribution, formulated as same-author retrieval and authorship verification, formulated as same-author binary decision. We benchmark 47 neural models and three non-neural baselines under a unified zero-shot protocol. Results show that authorship representation remains far from solved: the best retrieval model reaches only 0.258 Success@5, while the best verification model achieves 0.076 EER and 0.968 ROC-AUC. The leaderboard also reveals a meaningful task split, with different model families leading retrieval and verification, and large performance differences across languages, genres, and lengths. These findings position AuthBench not only as a new benchmark, but as a diagnostic resource for studying when and why authorship representations succeed or fail. We release AuthBench, its evaluation toolkit, and benchmark data at https://github.com/mao-code/AuthBench and https://huggingface.co/datasets/MaoXun/AuthBench.
Problem

Research questions and friction points this paper is trying to address.

authorship representation
multilingual benchmark
genres
document lengths
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

large-scale multilingual benchmark
authorship representation
zero-shot protocol
💼 Related Jobs
No related jobs found.
M
MaoXun Huang
Department of Computer Science, Cornell University
Zhenxing Zhang
Zhenxing Zhang
School of computing, Dublin City University
machine learningcomputer visioninformation retrieval
C
Claire Cardie
Department of Computer Science, Cornell University