🤖 AI Summary
This study addresses the challenge of tracing generative pipelines in AI-driven influence operations by constructing Propagia, the first French propaganda corpus. By integrating topic modeling, sentiment analysis, prompt leakage detection, and rewriting-based attribution techniques, this work reverse-engineers the content generation process. The research reveals distinct stylistic characteristics of AI-generated propaganda, identifies evidence of prompt leakage across 50 websites, and successfully attributes generated content to Llama-3 and Mistral models. Collectively, these findings establish a systematic methodological framework for the forensic analysis and provenance identification of AI-generated content, offering critical insights into detecting and mitigating automated disinformation campaigns.
📝 Abstract
We present a forensic analysis of the generation pipeline behind a recent AI-driven influence campaign. We introduce PROPAGIA, a corpus of 2,646 propagandist French articles from the Storm-1516/CopyCop campaign disclosed by VIGINUM and INSIKT GROUP in 2025. For comparison, we rely on SIPA, a corpus of human-written French mainstream press from the same period. Using topic modeling, vagueness and sentiment analysis, we first isolate persuasion techniques characteristic of propaganda, with PROPAGIA far exceeding SIPA in vagueness, subjectivity and negativity, and citing fewer sources. We then find prompt instruction leaks on 50 of the 84 PROPAGIA websites, including a verbatim ten-point editorial specification accounting for several of these differences, together with high cross-article redundancy. Finally, we show that rewriting-based detection supports INSIKT GROUP's attribution to the Llama 3 family, but also suggests the involvement of Mistral-family models.