MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering model
为解决网络安全分析中的信息过载等问题,提出MITRE-SAGE模型,通过多代理检索增强生成框架整合语义和结构知识,提高基于LLM的问答系统可靠性。
为解决网络安全分析中的信息过载等问题,提出MITRE-SAGE模型,通过多代理检索增强生成框架整合语义和结构知识,提高基于LLM的问答系统可靠性。
This work proposes a novel lossless compression method based on the T5 architecture that effectively leverages structural redundancy in complex data while preserving the discrete symbolic nature of the original input. Unlike conventional compression techniques, which struggle to exploit such redundancy, and existing deep learning approaches that rely on continuous vector representations—thereby disrupting the discrete token structure—our method integrates discrete latent representations with off-policy reinforcement learning. Trained end-to-end, it directly optimizes the length of discrete symbol sequences without requiring external knowledge. By maintaining the integrity of the original token structure and semantics, the approach achieves superior compression ratios and enhanced generalization across diverse data types, significantly outperforming traditional compression algorithms.
This work investigates the quantum acceleration potential and practical feasibility of Grover’s algorithm for preimage attacks against 3-round Keccak-256. Method: Leveraging hardware-aware, end-to-end quantum resource modeling—including surface-code error correction—we quantify physical implementation overheads for the first time: ~3.2 million physical qubits, ultra-deep circuits, and severe error accumulation. Contribution/Results: While theoretical speedup reduces search complexity from $2^{57.8}$ to $2^{28.9}$, it is fully offset by prohibitive hardware costs; estimated attack runtime spans 43 days to 2,365 years—highly sensitive to device assumptions yet fundamentally bottlenecked by qubit count and circuit depth. The study confirms SHA-3’s resilience against practical quantum preimage attacks in the foreseeable future. Its key innovations include the first complete quantum circuit synthesis for Keccak-256, Toffoli gate optimization, and a unified logical–physical qubit complexity analysis framework.
To address the scarcity of high-quality opinion role labeling (ORL) data for opinion mining in low-resource settings, this paper proposes a novel ORL data construction method grounded in semantic role labeling (SRL). We systematically map PropBank-style SRL annotations onto a Holder-Expression-Target ternary schema—the first such effort—and introduce three principled strategies: syntactic tree pointer conversion, discontinuous constituent handling, and semantic consistency filtering, enabling reproducible extraction of ORL instances from OntoNotes. The resulting dataset comprises 97,169 high-quality predicate-argument pairs. Empirical evaluation demonstrates substantial improvements in opinion role identification under low-resource conditions, particularly in cross-task transfer learning and multi-task learning settings. This work establishes a scalable, linguistically grounded data foundation for knowledge transfer across opinion analysis tasks.
To address the scalability limitations, high communication overhead, and elevated privacy risks inherent in conventional centralized federated learning (FL), this paper proposes a hierarchical multi-edge FL framework tailored for large-scale, privacy-sensitive healthcare applications. The method introduces: (1) an adaptive client scoring mechanism integrating utility, energy efficiency, and data sensitivity; (2) an end-edge collaborative privacy-preserving architecture combining homomorphic encryption, differential privacy, and secure aggregation; and (3) fair and efficient model aggregation across multiple edge servers. Experiments on the eICU dataset demonstrate that the proposed approach significantly outperforms FedAvg, FedProx, and FedSelect in prediction accuracy and cross-regional fairness, while simultaneously reducing communication costs and energy consumption.
为解决网络安全分析中的信息过载等问题,提出MITRE-SAGE模型,通过多代理检索增强生成框架整合语义和结构知识,提高基于LLM的问答系统可靠性。
This work proposes a novel lossless compression method based on the T5 architecture that effectively leverages structural redundancy in complex data while preserving the discrete symbolic nature of the original input. Unlike conventional compression techniques, which struggle to exploit such redundancy, and existing deep learning approaches that rely on continuous vector representations—thereby disrupting the discrete token structure—our method integrates discrete latent representations with off-policy reinforcement learning. Trained end-to-end, it directly optimizes the length of discrete symbol sequences without requiring external knowledge. By maintaining the integrity of the original token structure and semantics, the approach achieves superior compression ratios and enhanced generalization across diverse data types, significantly outperforming traditional compression algorithms.
This work investigates the quantum acceleration potential and practical feasibility of Grover’s algorithm for preimage attacks against 3-round Keccak-256. Method: Leveraging hardware-aware, end-to-end quantum resource modeling—including surface-code error correction—we quantify physical implementation overheads for the first time: ~3.2 million physical qubits, ultra-deep circuits, and severe error accumulation. Contribution/Results: While theoretical speedup reduces search complexity from $2^{57.8}$ to $2^{28.9}$, it is fully offset by prohibitive hardware costs; estimated attack runtime spans 43 days to 2,365 years—highly sensitive to device assumptions yet fundamentally bottlenecked by qubit count and circuit depth. The study confirms SHA-3’s resilience against practical quantum preimage attacks in the foreseeable future. Its key innovations include the first complete quantum circuit synthesis for Keccak-256, Toffoli gate optimization, and a unified logical–physical qubit complexity analysis framework.
To address the scarcity of high-quality opinion role labeling (ORL) data for opinion mining in low-resource settings, this paper proposes a novel ORL data construction method grounded in semantic role labeling (SRL). We systematically map PropBank-style SRL annotations onto a Holder-Expression-Target ternary schema—the first such effort—and introduce three principled strategies: syntactic tree pointer conversion, discontinuous constituent handling, and semantic consistency filtering, enabling reproducible extraction of ORL instances from OntoNotes. The resulting dataset comprises 97,169 high-quality predicate-argument pairs. Empirical evaluation demonstrates substantial improvements in opinion role identification under low-resource conditions, particularly in cross-task transfer learning and multi-task learning settings. This work establishes a scalable, linguistically grounded data foundation for knowledge transfer across opinion analysis tasks.
To address the scalability limitations, high communication overhead, and elevated privacy risks inherent in conventional centralized federated learning (FL), this paper proposes a hierarchical multi-edge FL framework tailored for large-scale, privacy-sensitive healthcare applications. The method introduces: (1) an adaptive client scoring mechanism integrating utility, energy efficiency, and data sensitivity; (2) an end-edge collaborative privacy-preserving architecture combining homomorphic encryption, differential privacy, and secure aggregation; and (3) fair and efficient model aggregation across multiple edge servers. Experiments on the eICU dataset demonstrate that the proposed approach significantly outperforms FedAvg, FedProx, and FedSelect in prediction accuracy and cross-regional fairness, while simultaneously reducing communication costs and energy consumption.