🤖 AI Summary
研究通过区分通用与非通用重复内容,使用LLM辅助标注和监督分类器处理187,000条推文,以提高信息操作中基于重复内容的协调行为分析准确性。
📝 Abstract
Duplicate content is widely used to study coordinated behavior in social media information operations (IOs), but not all repetition provides equally meaningful evidence of coordination. Generic, reusable, or low-information posts may create noisy account-account links when projected into coordination graphs. We study this problem using 187,000 English-language tweets from six Russian Twitter Information Operations datasets. We introduce a generic/non-generic distinction for duplicate campaigns, label tweets using an LLM-assisted protocol with independent human validation, and train supervised classifiers over sentence embeddings to scale the labels. We construct duplicate campaigns using lexical similarity and two embedding-based methods. Generic campaigns are rare under lexical matching but account for nearly 39% of campaigns detected by embedding-based methods. Restricting graphs to non-generic campaigns reduces graph size and the largest connected component while increasing density, suggesting a smaller but more focused coordination structure. These findings show that duplicate-based coordination analysis should consider both textual similarity and semantic specificity.