Concept Navigation and Classification via Open Source Large Language Model Processing

📅 2025-02-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the low accuracy and poor interpretability of automated identification of latent concepts—such as frames, narratives, and themes—in textual data. To this end, we propose a human-in-the-loop concept discovery framework. Methodologically, we introduce the first open-source large language model (LLM)-driven concept navigation paradigm, integrating iterative prompt-based sampling, cross-domain text embedding and clustering, and an expert validation feedback loop—thereby tightly coupling automated summarization with human-in-the-loop verification. Experiments on AI policy debates, cryptocurrency news, and the 20 Newsgroups dataset demonstrate substantial improvements in political discourse analysis, media frame detection, and fine-grained topic classification. Our approach achieves a superior trade-off between accuracy and interpretability, offering a robust, transparent, and reproducible pathway for concept modeling in computational social science.

Technology Category

Application Category

📝 Abstract
This paper presents a novel methodological framework for detecting and classifying latent constructs, including frames, narratives, and topics, from textual data using Open-Source Large Language Models (LLMs). The proposed hybrid approach combines automated summarization with human-in-the-loop validation to enhance the accuracy and interpretability of construct identification. By employing iterative sampling coupled with expert refinement, the framework guarantees methodological robustness and ensures conceptual precision. Applied to diverse data sets, including AI policy debates, newspaper articles on encryption, and the 20 Newsgroups data set, this approach demonstrates its versatility in systematically analyzing complex political discourses, media framing, and topic classification tasks.
Problem

Research questions and friction points this paper is trying to address.

Detecting latent constructs in text
Classifying frames, narratives, topics
Enhancing accuracy with human validation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Open-Source LLMs for text analysis
Hybrid summarization with human validation
Iterative sampling for conceptual precision
💼 Related Jobs
No related jobs found.
M
M. Kubli
Department of Political Science, University of Zurich