Challenges in annotations by humans and LLMs: A case study of evaluative language

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenging and highly subjective task of annotating evaluative language by focusing on the Attitude subsystem within Appraisal Theory, specifically examining expressions of affect, judgment, and appreciation in TED Talks. Through a systematic comparison of annotation behaviors among trained linguists, novice linguists, and large language models (LLMs), it reveals for the first time the performance disparities between human expertise and LLMs in such nuanced tasks. By integrating three prompt engineering strategies with model fine-tuning, the research significantly enhances the LLM’s capacity for automatic classification of Appraisal categories, achieving an F1 score of 0.77 post-fine-tuning—surpassing novice linguists and approaching the performance of expert linguists. These findings demonstrate the promising potential of LLMs as effective assistants in complex theoretical annotation tasks within digital humanities.
📝 Abstract
In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.
Problem

Research questions and friction points this paper is trying to address.

evaluative language
annotation challenges
Appraisal theory
large language models
subjectivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Appraisal theory
large language models
evaluative language
annotation agreement
prompt engineering
💼 Related Jobs
No related jobs found.
M
Mirela Imamovic
University of Hildesheim, Institute of Translation Studies and Specialised Communication, Germany; University of Paris 8 Vincennes-Saint Denis, TransCrit, France
A
Aenne Cecilia Kristine Knierim
University of Hildesheim, Institute of Information Science and Language Technology, Germany
K
Khushi Pitroda
University of Hildesheim, Institute of Translation Studies and Specialised Communication, Germany
E
Ekaterina Lapshinova-Koltunski
University of Hildesheim, Institute of Translation Studies and Specialised Communication, Germany