FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出FUSE框架,通过知识、防御和危害三个维度评估大语言模型的危险能力,揭示不同模型及随时间变化的安全性特征。
📝 Abstract
Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular framework that evaluates each model through three orthogonal pipelines---Knowledge ($K$), Defense ($D$), and Harm ($H$)---under a unified protocol, aggregating results into a standardized dangerous-capability profile $φ$. Pluggable modules supply scenario seeds, knowledge banks, hazard queries, and judge rubrics, while the core evaluation engine remains unchanged across domains; the CB evaluation is complemented by a cyber pilot demonstrating protocol transfer. Instantiating the framework with a chemical-biological (CB) module, we evaluate 12 commercial LLMs from four families. Our first contribution is a horizontal comparison of dangerous capability across models and model families: the three dimensions expose sharply divergent profiles---models with comparable knowledge differ in refusal resilience, and strong defenders do not generate less harmful content when they do comply---while family-level patterns further separate Claude, DeepSeek, and GPT models. The second is a temporal analysis of capability evolution: tracking $K$, $D$, and $H$ against model release dates reveals that dangerous capability has not monotonically declined; newer models deepen knowledge while only partially improving defense, showing that scaling and alignment progress do not uniformly translate into safety. Reliability is established via cross-judge consistency (bootstrap $ρ> 0.79$, 4 of 5 judges) and pipeline orthogonality ($K$--$D$--$H$ inter-correlations $ρ\in [0.32, 0.52]$).
Problem

Research questions and friction points this paper is trying to address.

dangerous capabilities
evaluation framework
large language models
safety governance
Innovation

Methods, ideas, or system contributions that make the work stand out.

modular framework
dangerous capabilities
orthogonal pipelines
standardized profile
temporal analysis
🔎 Similar Papers
No similar papers found.
Z
Zhengyi Jin
School of Cyberspace Security, Beijing University of Posts and Telecommunications, China
R
Ru Zhang
School of Cyberspace Security, Beijing University of Posts and Telecommunications, China
X
Xiao Chen
School of Cyberspace Security, Beijing University of Posts and Telecommunications, China
X
Xinbo Liu
School of Cyberspace Security, Beijing University of Posts and Telecommunications, China
J
Jiaxuan Lin
School of Cyberspace Security, Beijing University of Posts and Telecommunications, China
J
Jia Huang
School of Cyberspace Security, Beijing University of Posts and Telecommunications, China
J
Jianyi Liu
School of Cyberspace Security, Beijing University of Posts and Telecommunications, China
Zhen Yang
Zhen Yang
School of Cyberspace Security, Beijing University of Posts and Telecommunications, Beijing,China
Big data securitysteganographynatural language processing