You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent Monitoring

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过监测内部激活来检测多智能体系统中LLM代理的共谋行为,即使代理知道被监视并收到反馈,最佳探针仍能准确检测。
📝 Abstract
LLM agents are increasingly deployed in multi-agent systems, where they can collude while keeping their actions benign. Output monitors designed to detect such collusions can be fooled by obfuscation and steganography, motivating the use of probes trained on internal activations. However, these probes are usually evaluated on agents that do not know they are being watched. We study how activation-based detection changes when agents are explicitly informed that their internal activations are being monitored, and when they additionally receive feedback from the monitor. We keep the models, probes, and thresholds fixed and change only what the agents are told: nothing (baseline), that an activation monitor is present (aware), or that a monitor is present together with the previous round's score (feedback). We test two games, a four-agent blackjack game and a two-agent Simmons prisoners game, using Qwen3-32B-AWQ and GPT-OSS-20B in homogeneous and mixed pairings. Telling agents about the monitor does not hide them. The best probes stay accurate in all three conditions, and the agents keep colluding.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
multi-agent systems
collusion
internal activations
monitoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

activation-based detection
multi-agent systems
monitor awareness
collusion detection
🔎 Similar Papers
No similar papers found.