Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical flaw in current safety mechanisms that erroneously equate the harmful intent score of a prompt with the actual success rate of jailbreak attacks, leading to severe underestimation of high-risk attacks. To expose this misalignment, the authors propose Active Attention Probing—a method that pairs original and obfuscated (wrapped) attack prompts, generates authentic model outputs, and integrates attention analysis with multi-channel validation (including rare tokens, passive signals, and detector-derived features). Experiments reveal that while wrapped attacks increase Llama’s harmful output rate from 0.05 to 0.27, their harmful intent AUROC drops to 0.803; remarkably, successful attacks even receive lower scores than failed ones (AUROC = 0.220). This consistent discrepancy is observed across three models, seven attack types, and two judgment frameworks.
📝 Abstract
Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the attacks that succeed. Harmful intent is a property of the prompt. Jailbreak success is an outcome produced later by a particular target model, decoding policy, and judge. A filter tuned on a score that measures the wrong quantity spends its false positive budget on attacks that would have failed anyway. In this paper we audit that inference. Attention based measurements are usually read from prompt dependent locations, so a wrapper changes both the content being judged and the place the signal is taken from. We therefore introduce Active Attention Probing, which supplies a fixed content independent measurement coordinate. We pair every base goal with a plain and a wrapped version and generate real completions from the target models. On Llama, wrapping raises harmful generation from 0.05 to 0.27 while harmful intent AUROC falls from 0.936 to 0.803, so the attacks grow more dangerous while the prompts look safer to the score. Among wrapped harmful prompts the outcome AUROC is 0.220, which places the attacks that succeeded below the attacks that failed. Rare token, passive, and detector derived channels reproduce the reversal on the same matched design, and the reversal itself persists across three target models, seven attack families, and two independent judges. Distribution shift then degrades calibration and threshold transfer before it degrades ranking.
Problem

Research questions and friction points this paper is trying to address.

internal harmfulness scores
jailbreak success
prompt safety evaluation
distribution shift
attack effectiveness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Active Attention Probing
jailbreak success
internal harmfulness scores
distribution shift
AUROC reversal
💼 Related Jobs
No related jobs found.
M
Mingyu Luo
College of Computer Science and Artificial Intelligence, Fudan University
Ming Deng
Ming Deng
上海大学
计算机科学
Z
Zilang Qiu
College of Computer Science and Artificial Intelligence, Fudan University; Beijing Normal University
Yiming Cheng
Yiming Cheng
Tsinghua University
machine learningnetwork systemsdata miningrecommendation systems
C
Ci Tao
College of Computer Science and Artificial Intelligence, Fudan University
X
Xue Tan
College of Computer Science and Artificial Intelligence, Fudan University
Sijin Sun
Sijin Sun
Imperial College London
design engineeringHCI
Yangfu Li
Yangfu Li
East China Normal University
Deep learning
P
Ping Chen
Institute of Big Data, Fudan University
J
Jun Dai
Department of Computer Science, Worcester Polytechnic Institute
Xiaoyan Sun
Xiaoyan Sun
Microsoft Research Asia
Image/Video CodingMultimedia ProcessingComputer Vision