When Lexical Change Misleads: Rethinking Dynamic Topic Model Evaluation with Traditional and LLM-Based Metrics

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of traditional coherence metrics in evaluating dynamic topic models due to vocabulary shifts. We propose a vocabulary-shift-aware evaluation framework and, through human annotation and cross-domain experiments across multiple datasets, demonstrate the limitations of conventional metrics while advocating for LLM-based semantic measures as complementary signals. Results indicate that LLM semantic metrics align strongly with human judgment (ρ=0.721), significantly outperforming traditional approaches, and function complementarily rather than as replacements. This hierarchical evaluation strategy effectively enhances the reliability and robustness of dynamic topic model assessment, offering a novel perspective for evaluation paradigms in the field.
📝 Abstract
Dynamic topic models capture evolving word distributions, but traditional coherence metrics may fail when vocabulary changes while semantic meaning persists. We evaluate 120 topics from CoNTM and DLDA across NYT, DBLP, and arXiv, using three human annotators and Low, Medium, and High lexical-change categories. Traditional temporal coherence shows highly variable agreement with human judgments ($ρ$=-0.256 to 0.614). In contrast, LLM-based semantic similarity agrees strongly with human semantic judgments for CoNTM on NYT ($ρ$=0.609), DBLP ($ρ$=0.721), and arXiv ($ρ$=0.502), but is less consistent for DLDA. Lexical-change stratification reveals variation hidden by aggregate evaluation. We therefore advocate lexical-change-aware evaluation, jointly reporting traditional coherence and LLM-based semantic measures as complementary rather than interchangeable signals.
Problem

Research questions and friction points this paper is trying to address.

Dynamic Topic Models
Evaluation Metrics
Lexical Change
Temporal Coherence
LLM-based Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Topic Model Evaluation
Lexical-Change-Aware
LLM-Based Semantic Similarity
Temporal Coherence
Complementary Metrics