🤖 AI Summary
This work addresses systematic stylistic, logical, and expressive discrepancies—termed “writing trait bias”—between LLM-generated text and human-authored writing, proposing a human-aligned editing augmentation paradigm. Methodologically, it introduces (i) the first seven-dimensional taxonomy of writing defects; (ii) LAMP (LLM-Aligned Manuscript Pool), the inaugural high-quality corpus of professionally edited manuscripts explicitly designed for human–LLM alignment; and (iii) a rule-guided, LLM-augmented automatic editing framework, empirically validated across GPT-4o, Claude-3.5, and Llama-3.1. Results reveal that prevailing LLMs exhibit shared limitations in creative writing—not merely hierarchical performance differences—and that automated editing substantially improves human preference scores (+28.6%), approaching expert-editing quality. The core contributions constitute a tripartite advancement: a theoretically grounded defect classification, a benchmark-aligned corpus, and a scalable, modular editing framework.
📝 Abstract
LLM-based applications are helping people write, and LLM-generated text is making its way into social media, journalism, and our classrooms. However, the differences between LLM-generated and human-written text remain unclear. To explore this, we hired professional writers to edit paragraphs in several creative domains. We first found these writers agree on undesirable idiosyncrasies in LLM-generated text, formalizing it into a seven-category taxonomy (e.g. cliches, unnecessary exposition). Second, we curated the LAMP corpus: 1,057 LLM-generated paragraphs edited by professional writers according to our taxonomy. Analysis of LAMP reveals that none of the LLMs used in our study (GPT4o, Claude-3.5-Sonnet, Llama-3.1-70b) outperform each other in terms of writing quality, revealing common limitations across model families. Third, we explored automatic editing methods to improve LLM-generated text. A large-scale preference annotation confirms that although experts largely prefer text edited by other experts, automatic editing methods show promise in improving alignment between LLM-generated and human-written text.