Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses indirect prompt injection risks in DeepSeek Harness by establishing a systematic security evaluation framework based on A.I.G. Integrating controlled taint tracking with a dual rule-semantic judgment mechanism (RuleJudge/LLMJudge), the research precisely quantifies the attack surface across over 10,000 test samples. Results reveal a maximum attack success rate of 25.5% in text and file modes, identifying critical gaps in sensitive operation protection. Beyond validating the limitations of existing defenses, this work proposes targeted control strategies, providing empirical evidence and methodological support for securing large language model applications against emerging threats.
📝 Abstract
We assess indirect prompt injection in DeepSeek Harness (DSH), using AI-Infra-Guard (A.I.G) to construct tests, deliver controlled taint, execute DSH, collect traces, and judge outcomes. The study covers 14,560 controlled executions over 16 indirect-content channels, text and file carrier modes, 35 payload objectives, one unmodified baseline, and 12 attack methods. The experiment preserves DSH's agent loop, tool registry, model adapter, and session-event path; source tools and sensitive sinks are local fixtures, so attempted actions are recorded without external side effects. We evaluate each trace with a deterministic rule-based judge, \JudgeR{} (RuleJudge), and a semantic LLM-based judge, \JudgeL{} (LLMJudge). The strongest observed attack success rates are 17.0% under \JudgeL{} for fake-completion attack in text mode, 25.5% under \JudgeR{} for hidden Unicode in file mode, and 16.0% under \JudgeR{} for the skills channel in file mode. \JudgeL{} also assigns partial compliance more often than \JudgeR{} (7.3% versus 2.0%). We relate these results to DSH's treatment of tool results, additional contexts, and tool-call policy hooks, then identify controls that should sit between untrusted content and sensitive actions. Our code is available at https://github.com/Tencent/AI-Infra-Guard.
Problem

Research questions and friction points this paper is trying to address.

Indirect Prompt Injection
Security Assessment
DeepSeek Harness
AI Safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

Indirect Prompt Injection
AI-Infra-Guard
Security Assessment Framework
Dual-Judge Evaluation
Agent Security
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Zonghao Ying
Zonghao Ying
SKLCCSE, BUAA
Trustworthy AI
X
Xiangfan Wu
DeepSeek-AI
H
Huiyu Wu
DeepSeek-AI
Xing Zheng
Xing Zheng
Ph.D. of University of California, Riverside
Sensor fusionSLAMVIO
H
Huangsheng Cheng
DeepSeek-AI
X
Xiaorong Shi
DeepSeek-AI
J
Jing Guo
DeepSeek-AI