Challenges in annotations by humans and LLMs: A case study of evaluative language
This paper compares human and large language model (LLM) annotations of evaluative language in TED talk transcripts using Appraisal theory, finding that LLMs outperform both trained linguists and trainees in complex annotation tasks, thereby demonstrating their potential to support digital humanities research.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Technical Summary: Challenges in Annotations by Humans and LLMs: A Case Study of Evaluative Language
Problem Statement
The paper addresses the increasing complexity of annotation tasks in digital humanities and computational social sciences, specifically focusing on the identification and classification of evaluative language. While Large Language Models (LLMs) have shown promise in unsupervised NLP tasks, their ability to match or exceed human performance in highly subjective, complex theoretical frameworks remains unverified. The study focuses on Appraisal theory (Martin & White, 2005), a functional model for evaluative language, specifically its Attitude subsystem (Affect, Judgement, Appreciation). The authors highlight that manual annotation of these categories is time-consuming, prone to low inter-annotator agreement (IAA), and complicated by context-dependency, implicit meanings, and the difficulty of distinguishing between subtle categories. The core research problem is to determine if LLMs encounter the same challenges as human annotators in these complex tasks and whether they can serve as effective tools for resolving such annotation difficulties.
Methodology
The study employs a mixed-methods case study using a corpus of 190 English TED talk transcripts (the EmotionalizTED corpus) across eleven domains (e.g., Science, Politics, Medicine). The methodology is divided into three phases:
Human Annotation:
- Annotators: 24 linguistics students in training (non-native English speakers, mostly German) and one senior researcher (serving as the gold standard).
- Task: Sentence-level classification of evaluative content. Annotators first determined if a sentence was evaluative (binary) and then classified it into Attitude categories: Affect, Judgement, Appreciation, Ambiguous, or Uncertain. Multi-labeling was permitted.
- Evaluation: Inter-Annotator Agreement (IAA) was measured using Cohen's Kappa (for binary evaluativeness) and Krippendorff's alpha (for multi-class subclassification). A qualitative error analysis was conducted to identify sources of disagreement.
Automatic Annotation (LLM):
- Prompt Engineering: Three distinct prompt variations were designed and compared using the qwen3-30b-a3b-instruct-25075 model. These included few-shot and zero-shot approaches, with variations in context, persona, constraints, and the inclusion of Chain-of-Thought (CoT) reasoning.
- Model Selection & Tuning: The best-performing prompt was applied to three different LLMs. The most effective model was subsequently fine-tuned using QLoRA to optimize performance.
- Process: The LLMs followed a two-step process: (1) binary classification of evaluativeness, and (2) multiclass classification of Attitude categories for evaluative sentences.
Comparative Analysis:
- The performance of the fine-tuned LLM was compared against the gold standard (senior researcher) and the student annotators.
- The study specifically investigated whether LLMs struggled with the same specific linguistic phenomena (e.g., implicit meaning, target identification) that caused low agreement among humans.
Key Results
- Human Performance: Inter-annotator agreement among student linguists was variable and often low. Agreement on the binary "evaluative vs. non-evaluative" decision ranged from slight to moderate (Kappa 0.51 in some domains like Entertainment and Politics, but dropping significantly in History and Psychology). Agreement on the specific Attitude subclasses (Affect, Judgement, Appreciation) was even lower, with Krippendorff's alpha frequently falling below 0.50 and sometimes into negative values.
- Sources of Human Error: Qualitative analysis revealed that human disagreements stemmed from:
- Difficulty distinguishing between explicit and implicit meanings.
- Confusion regarding the "target of evaluation" (e.g., whether an emotion belongs to the speaker or a third party mentioned).
- Challenges with long, complex sentences containing nested evaluative meanings.
- Variations in English proficiency and domain-specific knowledge.
- LLM Performance: The fine-tuned LLM achieved an F1-score of 0.77 against the gold standard.
- Comparison: The LLM outperformed the student annotators and, notably, achieved higher agreement with the senior researcher (gold standard) than the students did.
- Challenge Alignment: The study found that LLMs did not necessarily struggle with the same specific issues as humans in the same ways; rather, they demonstrated a capacity to resolve complex classification tasks more consistently than the trained-but-inexperienced human cohort.
Key Contributions
- Annotation Scheme: The paper presents a specific annotation scheme for Appraisal theory's Attitude categories adapted for sentence-level analysis in spoken popular science discourse.
- Empirical Comparison: It provides a direct empirical comparison between human annotators (both trainees and experts) and LLMs in a complex, subjective linguistic task.
- Prompt Optimization: The study demonstrates the impact of prompt design (including CoT and few-shot strategies) on the performance of LLMs in pragmatic annotation.
Significance and Claims
The authors claim that LLMs can effectively aid in the resolution of complex annotation tasks, potentially surpassing the performance of linguists in training in terms of consistency and agreement with expert standards. The paper suggests that if LLMs can successfully annotate complex linguistic theories like Appraisal through prompting and fine-tuning, this opens new pathways for digital humanities research. Specifically, it implies that extensive, manually annotated corpora may not always be strictly necessary for classifying large portions of speech data, provided that LLMs are properly guided. The study concludes that LLMs offer a viable strategy for scaling annotation processes in interdisciplinary fields, though it acknowledges the need for task-by-task validation of model performance.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.