LLM-guided Semi-Supervised Approaches for Social Media Crisis Data Classification

📅 2026-05-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of efficiently classifying crisis-related social media posts in disaster management scenarios where labeled data are scarce. The authors propose and empirically evaluate two large language model (LLM)-guided semi-supervised learning approaches—LG-CoTrain and VerifyMatch—for low-resource crisis tweet classification. Experimental results demonstrate that LG-CoTrain significantly outperforms conventional methods, achieving the highest Macro F1 score with only 5–25 labeled examples per class, while VerifyMatch exhibits strong competitive performance and superior calibration. This work provides the first empirical validation of the effectiveness of LLM-guided mechanisms in this domain and shows that small models informed by LLMs can surpass zero-shot LLMs under specific conditions, substantially enhancing practical deployability.
📝 Abstract
Semi-supervised learning approaches have been investigated as a means to enhance the analysis of social media data in disaster management contexts. In this work, we present the first empirical evaluation of large language model (LLM) guided semi-supervised learning for crisis related tweet classification. We compare two recent LLM assisted semi-supervised methods, VerifyMatch and LLM guided Co-Training ( LG-CoTrain), against established semi-supervised baselines. Our results show that LG-CoTrain significantly outperforms classical semi-supervised approaches in low resource settings with 5, 10 and 25 labeled examples per class, achieving the highest averaged Macro F1 across events. VerifyMatch achieves competitive performance while also demonstrating strong calibration properties. As the number of labeled examples increases, the performance gap narrows and Self Training emerges as a strong baseline. We further observe that compact semi-supervised models can, in some cases, outperform very large LLMs operating in zero-shot settings. This finding highlights the potential of transferring knowledge from LLMs into smaller and more deployable models through LLM guided semi-supervised learning, offering a practical pathway for real world disaster response applications. Our project repository on Github is here.
Problem

Research questions and friction points this paper is trying to address.

semi-supervised learning
social media crisis classification
low-resource setting
large language models
disaster response
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-guided semi-supervised learning
crisis tweet classification
low-resource learning
knowledge transfer
model calibration
J
Jacob Ativo
Department of Computer Science, California State University, East Bay
B
Bharaneeshwar Balasubramaniyam
Department of Computer Science, Kansas State University
A
Anh Tran
Independent Researcher
K
Khushboo Gupta
Department of Computer Science, University of Illinois at Chicago
Hongmin Li
Hongmin Li
California State University, East Bay
Text ClassificationDomain AdaptationNLPMachine Learning
Doina Caragea
Doina Caragea
Kansas State University
deep learningtext miningdata miningdata science
Cornelia Caragea
Cornelia Caragea
University of Illinois at Chicago
Natural Language ProcessingDeep LearningInformation RetrievalArtificial Intelligence