You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建包含4735个诊断项目的基准,评估大语言模型在理解和解释中文社交媒体中含蓄和幽默评论的社会语用意义的能力。
📝 Abstract
Chinese online comments often convey social meaning through indirect and playful language that is hard to interpret without context. Existing evaluations largely organize items around predefined phenomena or controlled pragmatic categories, leaving open whether models can distinguish plausible readings of what a naturally occurring comment is doing in a particular exchange. We introduce a benchmark for evaluating whether LLMs can recover such situated pragmatic meanings. From more than 200,000 public Chinese social media interaction records, we construct 4,735 human-validated diagnostic items, each pairing a target comment with reconstructed preceding context and plausible misreadings. We evaluate eight LLMs as both question writers and solvers in a cross-writer setting. The task is challenging: the strongest model achieves 81.42% leave-writer-out accuracy. Across all eight models, the mean leave-writer-out accuracy is 68.70% while human accuracy was 90.8%. Case analysis shows that models often recognize broad irony or playfulness while misidentifying the mechanism or interactional move.
Problem

Research questions and friction points this paper is trying to address.

social pragmatic inference
indirect comments
playful language
Chinese online comments
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

social pragmatic inference
indirect and playful language
benchmarking
large language models
cross-writer setting
🔎 Similar Papers
No similar papers found.