Do Large Language Models Perform Well on Comprehending Poetic Logic in Modern Chinese Poetry?

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出Peony基准,通过四个任务评估大型语言模型理解现代中文诗歌独特逻辑的能力,揭示现有模型的局限。
📝 Abstract
Large Language Models (LLMs) have achieved significant progress across a wide range of natural language processing (NLP) tasks, yet their ability to understand literary texts, particularly modern Chinese poetry, remains largely unexplored. The unique literary characteristics of modern Chinese poetry necessitate a distinct form of reasoning for effective comprehension. Unlike conventional texts that convey clear information, the unique "poetic logic" of modern Chinese poetry requires a holistic reasoning approach that goes beyond superficial semantic analysis to be understood. However, current evaluation paradigms largely ignore this critical dimension. To address this gap, we propose Peony, the first benchmark specifically designed for evaluating the poetic logic of modern Chinese poetry. We define poetic logic as four tasks across three levels, namely stanza, line, and imagery, and systematically evaluate and analyze six mainstream LLMs based on Peony. We evaluate these models under both non-thinking and thinking configurations. The experimental results reveal the limitations of current LLMs in understanding the poetic logic of modern Chinese poetry and validate the effectiveness and necessity of Peony. Our data and code will be available.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Poetic Logic
Modern Chinese Poetry
Comprehension
Innovation

Methods, ideas, or system contributions that make the work stand out.

Poetic Logic
Modern Chinese Poetry
Benchmark
Large Language Models
Holistic Reasoning