Authorship Attribution in the Era of LLMs: Problems, Methodologies, and Challenges
The rise of large language models (LLMs) has intensified authorship attribution challenges—namely, distinguishing human-authored text from LLM-generated content and resolving ambiguous attribution in human-AI collaborative writing. To address this, we propose the first four-category authorship taxonomy for the LLM era: human-authored, LLM-generated, LLM-attributed, and human-AI collaborative. We systematically survey detection methodologies across four paradigms: statistical features, neural representations, attribution graphs, and prompt engineering—covering models including BERT, RoBERTa, and Llama, as well as watermarking and probability calibration techniques. Further, we introduce a unified evaluation framework balancing cross-domain generalizability and decision interpretability, and establish the field’s first dynamically updated resource repository (llm-authorship.github.io). Our work provides both theoretical foundations and a practical roadmap for enhancing detection accuracy and transparency in LLM-era authorship attribution.