ALERT: Zero-shot LLM Jailbreak Detection via Internal Discrepancy Amplification
This work addresses the challenge of detecting novel jailbreak attacks against large language models under zero-shot settings, where existing methods often fail to generalize. The authors propose a multi-granularity internal representation discrepancy amplification framework that systematically identifies security-critical components by analyzing model activations at the layer, module, and token levels. By integrating hierarchical, modular, and token-wise feature enhancement mechanisms and employing two complementary classifiers for joint decision-making, the approach effectively uncovers the model’s intrinsic discriminative safety signals. Evaluated on three mainstream safety benchmarks, the method consistently ranks among the top two, achieving average accuracy and F1 score improvements of 10%–40% over the strongest baselines—the first to demonstrate high-precision jailbreak detection in a zero-shot scenario.