Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建层级防御体系保护大型语言模型,但发现各层间存在依赖性,导致整体效果不如预期,需测量实际效果以优化防御策略。
📝 Abstract
Practitioners defend large language models (LLMs) by stacking defenses, assuming the layers compound. A stack is an ensemble, and ensembles compound only under a condition the LLM security literature recommends but never measures: the members must fail on different inputs. Two instruments make that measurable. The Adversary Access-Tier Model (AATM) grades an adversary by the access it holds, from system-only (A0) to influence over training data (A4). A cost model sorts defenses into five classes of inference-time overhead; because two classes require training weights or reading activations, they tier the defender as AATM tiers the adversary. From these we derive how a stack behaves, and the quantities a defender cares about diverge: coverage saturates within a tier, cost rises by class, false refusals accumulate as a union, and residual attack success falls multiplicatively only under independence. We measure that independence. Running one adaptive adversary against a seven-layer stack, failure correlation is positive in all fifteen measurable pairs ($φ$ from $0.30$ to $0.75$), and the joint residual exceeds the multiplicative prediction by up to $0.172$. Stratifying on behavior difficulty dissolves most of the association, so the dependence is predominantly common-cause, but it survives permutation inference, majority-vote grader labels, and externally calibrated thresholds. The same stack refuses four in five benign prompts while remaining statistically indistinguishable from its strongest single layer. The dependence is architectural rather than sampling-based: members correlate through the model they all wrap, so no wider member pool weakens it. Diversity therefore selects stack members but does not predict what an assembled stack delivers, which has to be measured end to end.
Problem

Research questions and friction points this paper is trying to address.

Layered LLM Defenses
Failure Correlation
Adversary Access-Tier Model
Inference Cost
Independence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversary Access-Tier Model
Failure Correlation
Layered Defenses
Ensemble Security
A
Abrar Alotaibi
Information and Computer Science Department, King Fahd University of Petroleum and Minerals, Dhahran, Saudi Arabia
M
Muhammad Shahid Jabbar
SDAIA-KFUPM Joint Research Center for Artificial Intelligence, King Fahd University of Petroleum and Minerals, Dhahran, Saudi Arabia
Sadam Al-Azani
Sadam Al-Azani
Research Scientist, SDAIA-KFUPM Joint Research Center for AI, KFUPM
Artificial IntelligenceArabic NLPMultimodal LearningVideo AnalyticsSocial Computing
Moataz Ahmed
Moataz Ahmed
King Fahd University of Petroleum & Minerals
Artificial Intelligence