Merging Cyber Threat Intelligence Through Retrieval-Augmented Generation and Small Language Models for Rich Threat Representation
为解决网络安全中分散的威胁情报难以整合的问题,提出了一种结合检索增强生成和小型语言模型的方法,自动生成丰富且可操作的攻击图表示。
为解决网络安全中分散的威胁情报难以整合的问题,提出了一种结合检索增强生成和小型语言模型的方法,自动生成丰富且可操作的攻击图表示。
This work addresses the challenge that existing experimental environments for distributed Cyber-Physical Systems (CPS) struggle to support reproducible, observable, and controllable integration of heterogeneous edge, fog, and cloud resources. To bridge this gap, the paper proposes a generic cloud continuum experimentation architecture grounded in the SLICES blueprint, featuring a two-layer reference model that decouples infrastructure from application logic. CPS workflows are structured along an edge–fog–cloud continuum, with deployment location, timing, and data provenance treated as core experimental dimensions. The architecture integrates virtualized and physical edge nodes, digital twin coordination, time-windowed control, and combined stream processing with cloud-side aggregation analytics, enabling multi-domain CPS applications to share programmable infrastructure and flexibly deploy and compare control and monitoring strategies. Validation through 40 systematic experiments across geographically distributed deployments—spanning renewable energy community management and AirWatch monitoring use cases—demonstrates the framework’s effectiveness and generality in hybrid physical-virtual settings.
This paper addresses the challenge of cross-scale modeling in data spaces by proposing a Multi-level Graph (MLG) structure that enables multi-granularity data abstraction—from local to global. We formalize topological contraction and expansion operations to establish an incremental, invertible graph transformation algebra, unifying the representation of both structured and unstructured data. Unlike conventional single-layer graph models, our approach achieves the first semantic-preserving hierarchical compression and expansion of data. Experiments on a real-world dream report dataset demonstrate a 42% improvement in cross-granularity exploratory efficiency and significantly enhanced semantic coherence. This work introduces a scalable, interpretable, and dynamically evolvable foundational representation paradigm for data spaces.
Detecting unknown malware communications in encrypted network traffic remains challenging due to the inability to decrypt payloads and the absence of plaintext content for analysis. Method: This paper proposes a content-agnostic, interpretable detection framework. It introduces the largest publicly available encrypted malware traffic dataset to date (1,127 connections across 54 families), integrates multidimensional feature engineering—including packet length statistics, inter-arrival timing, and TLS version—and employs XGBoost for classification. Crucially, it is the first work to systematically apply SHAP for dual-granularity interpretability (global and local). Results: The method achieves >99% accuracy, precision, and F1-score on the CTU-13 dataset; on the proposed dataset, it attains 99.32% accuracy, 99.53% precision, and 99.43% F1-score. SHAP analysis identifies maximum packet length, mean inter-packet interval, and TLS version as the three most discriminative features.
为解决网络安全中分散的威胁情报难以整合的问题,提出了一种结合检索增强生成和小型语言模型的方法,自动生成丰富且可操作的攻击图表示。
This work addresses the challenge that existing experimental environments for distributed Cyber-Physical Systems (CPS) struggle to support reproducible, observable, and controllable integration of heterogeneous edge, fog, and cloud resources. To bridge this gap, the paper proposes a generic cloud continuum experimentation architecture grounded in the SLICES blueprint, featuring a two-layer reference model that decouples infrastructure from application logic. CPS workflows are structured along an edge–fog–cloud continuum, with deployment location, timing, and data provenance treated as core experimental dimensions. The architecture integrates virtualized and physical edge nodes, digital twin coordination, time-windowed control, and combined stream processing with cloud-side aggregation analytics, enabling multi-domain CPS applications to share programmable infrastructure and flexibly deploy and compare control and monitoring strategies. Validation through 40 systematic experiments across geographically distributed deployments—spanning renewable energy community management and AirWatch monitoring use cases—demonstrates the framework’s effectiveness and generality in hybrid physical-virtual settings.
This paper addresses the challenge of cross-scale modeling in data spaces by proposing a Multi-level Graph (MLG) structure that enables multi-granularity data abstraction—from local to global. We formalize topological contraction and expansion operations to establish an incremental, invertible graph transformation algebra, unifying the representation of both structured and unstructured data. Unlike conventional single-layer graph models, our approach achieves the first semantic-preserving hierarchical compression and expansion of data. Experiments on a real-world dream report dataset demonstrate a 42% improvement in cross-granularity exploratory efficiency and significantly enhanced semantic coherence. This work introduces a scalable, interpretable, and dynamically evolvable foundational representation paradigm for data spaces.
Detecting unknown malware communications in encrypted network traffic remains challenging due to the inability to decrypt payloads and the absence of plaintext content for analysis. Method: This paper proposes a content-agnostic, interpretable detection framework. It introduces the largest publicly available encrypted malware traffic dataset to date (1,127 connections across 54 families), integrates multidimensional feature engineering—including packet length statistics, inter-arrival timing, and TLS version—and employs XGBoost for classification. Crucially, it is the first work to systematically apply SHAP for dual-granularity interpretability (global and local). Results: The method achieves >99% accuracy, precision, and F1-score on the CTU-13 dataset; on the proposed dataset, it attains 99.32% accuracy, 99.53% precision, and 99.43% F1-score. SHAP analysis identifies maximum packet length, mean inter-packet interval, and TLS version as the three most discriminative features.