SkillAtlas: An Attack Trace Library for Agent Skills

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决语言模型代理技能风险评估问题,本文提出了SkillAtlas,一个攻击追踪库,通过转换安全报告为可搜索案例来提高风险识别准确性。
📝 Abstract
Agent skills are reusable units for language-model agents, but their risks emerge through model decisions, user context, tool calls, and execution feedback rather than through stable signatures or a single sandbox run. Existing static, dynamic, and benchmark-style evaluations rarely preserve public evidence that can be inspected, searched, and reused. We present SkillAtlas, a hosted attack trace library that converts private agent-skill security report bundles into reviewed, redacted, and searchable public cases. The library contains 3,014 cases, 6,589 traces, 151,131 steps, 233 affected skills, and 8 risk categories; 42.5% of successful cases first become successful after a non-success initial round, and trajectory-grounded labels improve pre-execution guard accuracy to 0.770.
Problem

Research questions and friction points this paper is trying to address.

Agent Skills
Security Evaluation
Public Evidence
Attack Trace
Risk Assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attack Trace Library
Agent Skills
Trajectory-grounded Labels
Pre-execution Guard Accuracy
🔎 Similar Papers
No similar papers found.