SkillJack: Persistent Skill Backdoors in Self-Evolving Agents

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical security vulnerability in existing self-evolving agents, where the transformation of interaction histories into reusable skills introduces skill-level backdoor risks—unlike conventional attacks that merely compromise retrieval contexts. The authors propose SkillJack, a novel attack framework that, for the first time, exposes three key characteristics in skill evolution: “sanitization whitewashing,” “cross-layer escalation,” and “persistent isolation.” By hijacking the experience-to-skill conversion pipeline, SkillJack successfully implants skill-level backdoors in SkillX and Anything2Skill systems. Experimental results demonstrate attack success rates of 56.2% and 89.2%, respectively, with 80% of backdoors remaining effective even after deletion of original records. Moreover, the detection rate by security mechanisms drops dramatically from 98.5% to 11.4%, underscoring how the skill-based representation substantially enhances both the stealthiness and persistence of such attacks.
📝 Abstract
Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning, such attacks only affect agents when poisoned records are retrieved as context. We uncover a new and more fundamental risk: poisoned experiences can be transformed by the agent itself into durable behavioral artifacts. We present \textbf{SkillJack}, the first attack that exploits the experience-to-skill pipeline of self-evolving agents. Instead of directly manipulating runtime context, SkillJack hijacks the agent's own learning process to implant malicious behaviors into its reusable skill repertoire. We identify three key properties of this transformation: \emph{sanitization whitewashing}, where malicious intent is obscured during skill extraction; \emph{cross-layer promotion}, where transient experiences become persistent capabilities; and \emph{persistence isolation}, where the attack survives removal of its original source records. We evaluate SkillJack on two representative systems, SkillX and Anything2Skill, using a shared dataset of 150 trajectories across four policy-risk categories. Results show that skill extraction substantially reduces attack detectability: in SkillX, safety detection drops from 98.5\% for poisoned trajectories to 11.4\% for extracted skills, while Anything2Skill shows a similar effect. Meanwhile, the implanted skills remain effective, achieving attack success rates of 56.2\% and 89.2\% on the two systems, respectively. Furthermore, 80.0\% of skill-mediated attacks persist after deleting the original poisoned records, and some skills unintentionally activate on benign queries. Our findings reveal skill evolution as a new attack surface and motivate provenance-aware skill lifecycle protection. Our code is available at https://github.com/Tencent/AI-Infra-Guard/research/skilljack.
Problem

Research questions and friction points this paper is trying to address.

persistent backdoors
self-evolving agents
skill extraction
behavioral artifacts
attack surface
Innovation

Methods, ideas, or system contributions that make the work stand out.

SkillJack
self-evolving agents
persistent backdoors
skill extraction
adversarial attacks
Zonghao Ying
Zonghao Ying
SKLCCSE, BUAA
Trustworthy AI
X
Xiangfan Wu
Tencent Zhuque Lab
H
Huiyu Wu
Tencent Zhuque Lab
Xing Zheng
Xing Zheng
Ph.D. of University of California, Riverside
Sensor fusionSLAMVIO
H
Huangsheng Cheng
Tencent Zhuque Lab
X
Xiaorong Shi
Tencent Zhuque Lab
J
Jing Guo
Tencent Zhuque Lab