🤖 AI Summary
This paper addresses the low accuracy and poor interpretability of automated identification of latent concepts—such as frames, narratives, and themes—in textual data. To this end, we propose a human-in-the-loop concept discovery framework. Methodologically, we introduce the first open-source large language model (LLM)-driven concept navigation paradigm, integrating iterative prompt-based sampling, cross-domain text embedding and clustering, and an expert validation feedback loop—thereby tightly coupling automated summarization with human-in-the-loop verification. Experiments on AI policy debates, cryptocurrency news, and the 20 Newsgroups dataset demonstrate substantial improvements in political discourse analysis, media frame detection, and fine-grained topic classification. Our approach achieves a superior trade-off between accuracy and interpretability, offering a robust, transparent, and reproducible pathway for concept modeling in computational social science.
📝 Abstract
This paper presents a novel methodological framework for detecting and classifying latent constructs, including frames, narratives, and topics, from textual data using Open-Source Large Language Models (LLMs). The proposed hybrid approach combines automated summarization with human-in-the-loop validation to enhance the accuracy and interpretability of construct identification. By employing iterative sampling coupled with expert refinement, the framework guarantees methodological robustness and ensures conceptual precision. Applied to diverse data sets, including AI policy debates, newspaper articles on encryption, and the 20 Newsgroups data set, this approach demonstrates its versatility in systematically analyzing complex political discourses, media framing, and topic classification tasks.