MOLE: Detecting Insider Threats in AI Agents

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入MOLE基准测试,解决了AI代理在操作过程中可能泄露模型权重、污染训练数据或削弱发布门的问题,该方法能有效检测常规工作中的潜在威胁。
📝 Abstract
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives and agent refusal does not predict completion. MOLE enables comparison of 40 monitors across corpus generators, observability levels, and threats; even the best evaluated monitor in our single-day audit-event comparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improves budget-AUC by 10% over applying it to every account-day at comparable modeled cost.
Problem

Research questions and friction points this paper is trying to address.

Insider Threats
AI Agents
Model Misalignment
Prompt Injection
Data Poisoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

MOLE
internal threats
benchmark
AI agents
monitor performance
🔎 Similar Papers