From Repetition to Recognition: Inductive Discovery of Disinformation Narratives

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了虚假信息叙事标签自动生成的问题,通过无监督方法如聚类和图社区检测,并提出三层评估框架以优化生成效果。
📝 Abstract
In disinformation datasets, narratives are often understood as recurring interpretive patterns that group texts under narrative labels. Recent work formalized narrative mining as inductively inferring narrative labels from corpora, but its evaluation stays tied to predefined taxonomies, a closed-world setting that cannot capture narratives absent from the reference labels. We introduce a three-tier evaluation framework for unsupervised narrative label generation: recovery (against a corpus's own taxonomy), mining (against external label sets), and discovery (without predefined labels). Applying it, we compare clustering-based and graph-community-based pipelines across seven disinformation datasets, with human validation of discovery on two. The two families are complementary under automated metrics, but in a corpus with two prominent topics, clustering can reduce one topic to 2% of generated labels while graph-based pipelines stay balanced. Discovery validation also reveals many singletons (narrative labels derived from single claims, 30-62% of graph outputs), which clustering cannot produce. Annotators confirm many as recognizable disinformation narratives, suggesting that in open-world discovery the repetition assumed by narrative mining may be recognized outside the corpus, not within it. We release human-validated narrative candidate labels for the Climate Obstruction and PolyNarrative datasets to support taxonomy development and dataset extension.
Problem

Research questions and friction points this paper is trying to address.

disinformation
narrative mining
unsupervised narrative label generation
open-world discovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

unsupervised narrative label generation
three-tier evaluation framework
graph-community-based pipelines
clustering-based pipelines
discovery validation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Max Upravitelev
Technische Universität Berlin; German Research Center for Artificial Intelligence (DFKI); BIFOLD – Berlin Institute for the Foundations of Learning and Data; Centre for European Research in Trusted AI (CERTAIN); Johannes Gutenberg-Universität Mainz
Veronika Solopova
Veronika Solopova
Technische Universität Berlin
Computational linguisticsEthics of AI
Jing Yang
Jing Yang
Post-doc Researcher at the XplainNLP group, Quality and Usability lab at TU Berlin and BIFOLD
Natural Language ProcessingXAIFact-checkingMisinformation
C
Charlott Jakob
Technische Universität Berlin; German Research Center for Artificial Intelligence (DFKI); Johannes Gutenberg-Universität Mainz
A
Alexandra Tsiakalou
Technische Universität Berlin; German Research Center for Artificial Intelligence (DFKI)
N
Neda Foroutan
Technische Universität Berlin; German Research Center for Artificial Intelligence (DFKI)
Vera Schmitt
Vera Schmitt
Head of XplaiNLP Research Group at TU Berlin
NLP/LLMsXAIHCIDisinformationUsable Privacy