The Impact of Synthetic Data Augmentation on Discourse-Pragmatic Function Classification

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过使用Llama 3.1生成合成数据,并根据与真实数据的几何关系进行分类,以解决话语-语用功能分类中的数据稀疏问题,提高模型性能。
📝 Abstract
Synthetic data augmentation has become a common strategy for addressing class imbalance in NLP, but most approaches focus on the quantity and diversity of generated examples rather than their geometric relationship to real training data. We investigate this question in the context of discourse pragmatic function classification, a task where data sparsity is a structural feature rather than a collection artefact. Using 410 manually annotated instances of the English word look drawn from the British National Corpus, spanning four functions: Attention Signal, Directive, Discourse Marker, and Interjection. We generate synthetic training examples with Llama 3.1 and partition them by their cosine distance from real training data in RoBERTa embedding space. We compare six training conditions that differ in the placement of synthetic examples relative to the empirical decision boundary, while holding augmentation quantity constant across conditions. All augmented conditions improve macro F and accuracy over the real only baseline, but core proximal examples (NEAR) yield the largest gains in macro F (0.113), while a distance balanced mix achieves the highest accuracy (0.748). No condition improves AUC, indicating that augmentation shifts the decision boundary rather than improving the model's underlying probability estimates. These findings suggest that where synthetic examples land in representation space matters as much as how many are generated, with implications for low resource pragmatic classification more broadly.
Problem

Research questions and friction points this paper is trying to address.

Synthetic Data Augmentation
Discourse-Pragmatic Function Classification
Geometric Relationship
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic Data Augmentation
Discourse-Pragmatic Function Classification
Cosine Distance
RoBERTa Embedding Space
Data Sparsity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sara Sorahi
Institute of Linguistics, Faculty of Arts and Humanities, Heinrich Heine University Düsseldorf
Kevin Tang
Kevin Tang
University Professor, English Language and Linguistics, Heinrich-Heine-University Düsseldorf
computational linguisticslaboratory phonologypsycholinguisticslinguistic biomarkersprecision health
R
Reza Kazemian
Sun Yat-sen University, China