CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the emergent safety risks arising when individually safe skills are composed within autonomous AI agents. We propose Composkill, a novel framework accompanied by a dedicated benchmark to evaluate compositional safety. By constructing dual white-box and black-box attack systems to generate high-risk skill chains, this work reveals path-level properties of compositional risks and patterns of bridging reward decay. Experiments demonstrate that these attack strategies achieve risk chain formation rates of 83.3% and 80.6%, respectively. These findings confirm that existing single-skill certification mechanisms exhibit systematic deficiencies and limited interception capabilities in long-horizon tasks. Consequently, this research establishes a new paradigm for assessing the compositional safety of autonomous agents, highlighting critical vulnerabilities in current verification approaches when applied to complex, multi-step skill compositions.
📝 Abstract
Autonomous AI agents tackling Long Horizon Tasks depend on marketplace skills that are certified one at a time: a scanner returns a safety verdict for each skill and declares the ecosystem safe if every package passes. We show that this assumption fails under skill composition. A skill may pass the per-skill scanner individually yet participate in a risky composition when an agent connects its outputs, capabilities, or side effects with those of other scanner-passing skills. This makes skill composition risk a path level property rather than a node level property, explaining why existing skill scanners that inspect individual packages achieve limited interception. To study this threat, we present CompoSkill, a framework that constructs skill composition attacks through a dual attacker system. The white-box attacker knows the victim's installed skill pool and directly injects explicit skill-id sequences; the black-box attacker knows only a role profile, downloads the top marketplace skills for that scenario, builds a Skill Composition Graph, and searches for high risk chains whose implicit lures never name skill identifiers. We further construct CompoSkill-Bench, a benchmark of 1,140 records built from long-horizon professional workflows across five threats and six scenarios on OpenClaw and Nanobot. CompoSkill achieves risk Chain Formation Rates (CFR) up to 83.3% in the white box setting and 80.6% in the black box setting, while existing skill scanners block only a limited fraction of the risky compositions. Finally, we observe a bridge-bonus-then-hop-decay pattern: a bridge skill can increase attack success, but Attack Success Rate (ASR) decreases once additional hops make the risk chain longer than three skills. These results expose a systematic gap in single skill certification for autonomous AI agents.
Problem

Research questions and friction points this paper is trying to address.

Compositional Skill Chain Attacks
Autonomous AI Agents
Skill Composition Risk
Single Skill Certification
Long Horizon Tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Skill Chain Attacks
Skill Composition Graph
Dual Attacker System
Path-level Security Property
CompoSkill-Bench
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Mingxiao Liu
Hangzhou Dianzi University
Z
Zhoumian Jiang
Hangzhou Dianzi University
J
Jianan Ma
Hangzhou Dianzi University, Ant Group
Jian Zhang
Jian Zhang
Hangzhou Dianzi University
Dynamic networkGraph neural networkAnomaly detection
J
Jialuo Chen
Ant Group, Zhejiang University
X
Xinhao Deng
Ant Group, Tsinghua University
Z
Zhen Wang
Hangzhou Dianzi University