STINER: Automated Extraction of Strategic Cyber Threat Intelligence from X

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of extracting strategic-level Cyber Threat Intelligence (CTI) from informal social media text. We propose a fine-grained taxonomy tailored for strategic intelligence and construct the first expert-annotated threat alert corpus. Furthermore, we introduce a novel entity recognition framework integrating a DarkBERT domain-adaptive encoder with generative large language models. Experimental results demonstrate that this approach achieves an F1-score of 89.33%, significantly outperforming both general-purpose models and standalone LLMs. Notably, the system successfully provided early warnings for the SafePay ransomware campaign, validating its effectiveness and practical utility in efficiently extracting strategic CTI from complex, unstructured social media content.
📝 Abstract
Strategic Cyber Threat Intelligence (CTI) focuses on high-level insights, such as identifying targeted industries, attributing attacks to specific ransomware groups, and assessing the scale of data loss. Today, X (formerly Twitter) has become the fastest source for this intelligence, often hosting real-time breach announcements days before formal vendor reports. Converting this raw chatter into actionable intelligence requires navigating a complex linguistic landscape. Conventional Named Entity Recognition (NER) models struggle to parse the informal and highly irregular dialect of social media, creating a blind spot for automated defense systems. To address this challenge, we introduce STINER, a taxonomy and expert-annotated corpus for extracting strategic intelligence from social media streams. We construct a high-quality, expert-annotated dataset of 2,100 real-world alerts and propose a granular taxonomy of eight entity types centered on strategic pivots such as Threat Actor, Sector, and Location. We benchmark nine models across 12 evaluated configurations, spanning general-purpose and domain-adapted encoders, open-schema extraction, and generative LLMs in both zero-shot and fine-tuned settings. Domain-adapted encoders such as DarkBERT reach a strict F1-score of 89.33%, outperforming both general-purpose baselines and fine-tuned Large Language Models, which additionally incur substantially higher inference latency. Leveraging STINER-DarkBERT, we conduct a European threat landscape analysis for H1 2025. Our results align with official reporting on major targets while highlighting the distinct visibility profile of attacks in Spain, and illustrate how social-media-driven extraction can surface early signals of the SafePay ransomware campaign prior to its retrospective characterization in vendor threat landscape reports.
Problem

Research questions and friction points this paper is trying to address.

Strategic Cyber Threat Intelligence
Named Entity Recognition
Social Media
Informal Text Processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Strategic Cyber Threat Intelligence
Domain-adapted Encoders
Expert-annotated Corpus
Social Media NER
DarkBERT
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yasir Ech-Chammakhy
College of Computing, Mohammed VI Polytechnic University (UM6P), Ben Guerir, Morocco
O
Oussama Azrara
Deloitte Conseil, Paris, France
J
Jaafar Chbili
Deloitte Conseil, Paris, France
A
Anas Motii
College of Computing, Mohammed VI Polytechnic University (UM6P), Ben Guerir, Morocco