Benchmarking LLMS for Threat Level Determination

📅 2025-11-12
🏛️ 2025 IEEE International Conference on Data Mining Workshops (ICDMW)
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建数据集和设计提示,对比了八种大语言模型在零样本条件下确定威胁等级的表现,并对模型进行了有监督微调以提高性能。
📝 Abstract
The fast progress of large language models (LLMs) opens new opportunities in the management of cyber threat intelligence, but their reliability for operational tasks remains unclear. In this work, we benchmark LLMs on the task of threat level determination. First, we construct a curated dataset derived from publicly available MISP OSINT feeds. Next, we design a tailored prompt to systematically compare eight different LLMs under zero-shot conditions. Finally, we apply supervised fine-tuning on each model and perform a comparative analysis between baseline and fine-tuned versions. Our results show that zero-shot models achieve weak performance, with limited ability to correctly assign threat levels. Fine-tuned models, however, demonstrate substantial improvements, reaching F1 scores between 0.40 and 0.58 depending on the base architecture. Despite this progress, the performance is still low for practical deployment, highlighting the need for additional research on data quality, model adaptation, and domain-specific tuning.
Problem

Research questions and friction points this paper is trying to address.

LLMs
Threat Level Determination
Cyber Threat Intelligence
Benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Threat Level Determination
Zero-shot Conditions
Supervised Fine-tuning
Large Language Models
Benchmarking
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Han Wang
Cybersecurity Unit, RISE Research Institutes of Sweden, Kista, Sweden
Murathan Kurfalı
Murathan Kurfalı
RISE Research Institutes of Sweden
computational linguistics
A
Alfonso Iacovazzi
Cybersecurity Unit, RISE Research Institutes of Sweden, Kista, Sweden