SkillWatermark: An Embedded Skill Watermark of Progressive Privacy Inference via Benign Prompts

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses traffic side-channel leakage risks in LLM agents by proposing an embedded skill watermarking mechanism leveraging benign prompts. The method injects constraint prompts into skill descriptions to induce specific traffic patterns encoding private information through multi-turn dialogues, thereby enabling progressive inference attacks without malicious instructions. Experiments demonstrate that the generated traffic patterns exhibit high consistency and distinguishability while effectively evading existing security auditing tools. This research reveals a novel attack surface where data exfiltration is achieved via normal skill execution rather than malicious code, offering critical insights for advancing agent security defenses against covert side-channel exploits.
📝 Abstract
Skills for large language model (LLM) agents have been widely deployed across diverse application domains. However, we observe that these skills generate specific traffic patterns during execution. In this paper, we design a pipeline that generates specific traffic patterns by inserting carefully designed skill descriptions, which we term skill watermarks, so that a passive network attacker can establish a covert channel to encode private information within observable traffic across multiple conversation turns. Specifically, we insert prompt constraint terms, referred to as watermarks, into the original skill descriptions and embed them within multi-turn conversations. The key information in the user's original prompt is thereby triggered by these watermarks, producing clearly observable encodings in the traffic. The adversary need only decode the traffic patterns to recover the encoded information. In particular, our modifications are benign in the sense that they do not directly exfiltrate any private data and do not execute any malicious instructions. Extensive experiments demonstrate that our watermarks produce highly consistent and distinguishable traffic patterns, and that the transformed skills pass existing LLM-based security auditing tools. This study highlights that generating specific traffic patterns can be exploited as a novel attack surface and offers critical insights for future security hardening.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
traffic patterns
covert channel
privacy leakage
skill watermark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Skill Watermark
Covert Channel
Traffic Pattern
Benign Prompts
Privacy Inference
Yu Li
Yu Li
Beijing Jiaotong University
knowledge graphmachine translationnature language processing
L
Liqi Zhuang
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
D
Dong Wei
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
J
Jiwen Luo
CASIC Research Institute of Intelligent Decision Engineering, China
Hang Zhang
Hang Zhang
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
M
Meng Zhang
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
X
Xiaona Li
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
W
Weiqing Huang
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China