Seed-Guided Topic Discovery with Out-of-Vocabulary Seeds

📅 2022-05-04
🏛️ North American Chapter of the Association for Computational Linguistics
📈 Citations: 11
Influential: 1
📄 PDF
🤖 AI Summary
Existing seed-guided topic discovery methods rely on the closed-vocabulary assumption, rendering them incapable of handling out-of-vocabulary (OOV) user-provided seeds and failing to effectively incorporate semantic knowledge from pretrained language models (PLMs). To address this, we propose SeeTopic—the first framework extending seed-guided topic discovery to OOV scenarios. SeeTopic jointly models global semantic representations from PLMs and local contextual cues from the corpus to achieve semantic alignment and dynamic expansion of OOV seeds, and introduces a jointly optimized topic generation mechanism. Evaluated on three cross-domain real-world datasets, SeeTopic significantly improves topic coherence (+12.3%), accuracy (+9.7%), and diversity (+8.1%), while maintaining strong robustness under OOV seeds. This work establishes a more open and flexible paradigm for user-directed topic discovery.
📝 Abstract
Discovering latent topics from text corpora has been studied for decades. Many existing topic models adopt a fully unsupervised setting, and their discovered topics may not cater to users’ particular interests due to their inability of leveraging user guidance. Although there exist seed-guided topic discovery approaches that leverage user-provided seeds to discover topic-representative terms, they are less concerned with two factors: (1) the existence of out-of-vocabulary seeds and (2) the power of pre-trained language models (PLMs). In this paper, we generalize the task of seed-guided topic discovery to allow out-of-vocabulary seeds. We propose a novel framework, named SeeTopic, wherein the general knowledge of PLMs and the local semantics learned from the input corpus can mutually benefit each other. Experiments on three real datasets from different domains demonstrate the effectiveness of SeeTopic in terms of topic coherence, accuracy, and diversity.
Problem

Research questions and friction points this paper is trying to address.

Addresses out-of-vocabulary seeds in topic discovery.
Utilizes pre-trained language models for enhanced topic accuracy.
Improves topic coherence, accuracy, and diversity with SeeTopic.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Utilizes pre-trained language models
Handles out-of-vocabulary seeds
Combines general and local semantics
🔎 Similar Papers
2024-06-13arXiv.orgCitations: 0
University of Illinois at Urbana-Champaign | University of Washington
Y
Yu Zhang
University of Illinois at Urbana-Champaign, IL, USA
Y
Yu Meng
University of Illinois at Urbana-Champaign, IL, USA
X
Xuan Wang
University of Illinois at Urbana-Champaign, IL, USA
S
Sheng Wang
University of Washington, Seattle, WA, USA
Jiawei Han
Jiawei Han
Abel Bliss Professor of Computer Science, University of Illinois
data miningdatabase systemsdata warehousinginformation networks