AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对通过有限黑盒互动使弱代理获取强代理能力的问题,提出AgentLeak方法,利用执行差异识别关键行为并增强弱代理的能力。
📝 Abstract
Large language model (LLM) agents increasingly achieve long-horizon tasks by combining foundation models with explicit skills and implicit procedural knowledge acquired through execution. The resulting task-solving capabilities have become valuable proprietary assets, raising a new security question: can a substantially weaker attacker-controlled agent acquire the capabilities of a stronger proprietary agent through limited black-box interaction? Existing skill-stealing attacks recover explicit skill artifacts, yet we show that artifact leakage does not necessarily transfer capability: a weaker agent may possess the same skills but still fail because it lacks procedural behaviors implicitly realized by the stronger agent. Our key insight is that the skill execution gap itself forms a leakage surface, where missing behaviors are exposed through observable differences between successful victim executions and failed attacker executions. Based on this, we present AgentLeak, a black-box capability-cloning attack that identifies capability-critical behaviors from these execution differences and incorporates them into attacker-side skills, while keeping the attacker's model, harness, and tools unchanged. Across 20 task scenarios comprising 600 instances, diverse agent systems, and multiple backbone models, AgentLeak improves task pass rates by over 40% compared with direct skill reuse and recovers more than 80% of the victim--attacker capability gap. Our findings reveal a confidentiality risk in LLM agents: protecting explicit artifacts alone is insufficient, as observable execution behavior can leak the procedural knowledge required to reconstruct proprietary task-solving capabilities in low-capability and attacker-controlled agents.
Problem

Research questions and friction points this paper is trying to address.

Large language model (LLM) agents
black-box interaction
procedural knowledge
skill-stealing attacks
capability-cloning
Innovation

Methods, ideas, or system contributions that make the work stand out.

black-box capability-cloning
execution differences
procedural knowledge leakage
skill execution gap
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Xiaoting Lyu
Xiaoting Lyu
Beijing Jiaotong University
Y
Yuhong Wu
Xi’an Jiaotong University
Y
Yufei Han
INRIA
S
Shichang Liu
Xi’an Jiaotong University
L
Liang Zhang
University of Warwick
B
Bin Wang
Xi’an Jiaotong University
B
Bin Wang
Zhejiang Key Laboratory of Artificial Intelligence of Things (AIoT) Network and Data Security
Xiaobo Ma
Xiaobo Ma
University of Arizona
TransportationMachine LearningStatistics
W
Wei Wang
Xi’an Jiaotong University