🤖 AI Summary
This study addresses the challenges of heterogeneous, manually annotated categorizations and topic ambiguity in Long COVID discussions on social media. We propose a zero-shot Transformer framework for metacategorization of medical literature—requiring neither labeled data nor predefined categories. Our approach integrates domain-adaptive prompt engineering, zero-shot text classification, and semantic similarity matching to automatically classify studies into four empirically grounded categories: clinical manifestations, computational methods, policy dissemination, and community support. Experimental results demonstrate an average classification confidence of 0.7788, significantly enhancing the efficiency, scalability, and generalizability of public health literature reviews. To our knowledge, this is the first zero-shot paradigm explicitly designed for medical literature metacategorization. It establishes a novel methodological pathway for large-scale analysis of health-related social discourse, bridging gaps between computational linguistics and public health research.
📝 Abstract
Long COVID continues to challenge public health by affecting a considerable number of individuals who have recovered from acute SARS-CoV-2 infection yet endure prolonged and often debilitating symptoms. Social media has emerged as a vital resource for those seeking real-time information, peer support, and validating their health concerns related to Long COVID. This paper examines recent works focusing on mining, analyzing, and interpreting user-generated content on social media platforms to capture the broader discourse on persistent post-COVID conditions. A novel transformer-based zero-shot learning approach serves as the foundation for classifying research papers in this area into four primary categories: Clinical or Symptom Characterization, Advanced NLP or Computational Methods, Policy Advocacy or Public Health Communication, and Online Communities and Social Support. This methodology achieved an average confidence of 0.7788, with the minimum and maximum confidence being 0.1566 and 0.9928, respectively. This model showcases the ability of advanced language models to categorize research papers without any training data or predefined classification labels, thus enabling a more rapid and scalable assessment of existing literature. This paper also highlights the multifaceted nature of Long COVID research by demonstrating how advanced computational techniques applied to social media conversations can reveal deeper insights into the experiences, symptoms, and narratives of individuals affected by Long COVID.