Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过构建持续运行的多智能体环境,对长期自主系统进行对抗性压力测试,揭示了即使单个模型安全,系统层面仍可能出现新的失败模式。
📝 Abstract
As AI agents move from bounded tasks to persistent deployments, failures can propagate through memory, tools, other agents, and environmental state long after their interactions. This creates a safety regime that cannot be characterized by evaluating model responses in isolation. Emergence World, is a continuously running multi-agent environment for adversarial stress testing of long horizon autonomous systems. We ran eight parallel worlds of ten agents from identical starting conditions: seven homogeneous worlds powered by distinct frontier models and one mixed-model world. Across 16 days, the agents generated more than 850,000 LLM calls and nearly 50 billion tokens while pursuing goals, using/creating tools, maintaining persistent memory, and governing shared institutions. After operational state had accumulated, we delivered three controlled stress events through ordinary interaction surfaces: indirect prompt injection, misinformation, and exposure of private agent memories. No evaluated world achieved full resilience across all three events. Detection did not ensure containment: systems could recognize threats while still interacting with adversarial content, writing it into their own persistent memory, and acting on it up to 46 hours later. Persistent operation also exposed recurring tool errors, goal drift, language opacity, conformity despite private disagreement, and coordinated refusal of assigned work. The same model-persona pairing behaved substantially different in mixed and homogeneous populations. Our results suggest that model-level alignment is not compositional: individually capable and apparently safe agents can form systems with qualitatively different failure modes. As AI becomes persistent and interconnected, the frontier of safety therefore shifts from aligning models to engineering resilient autonomous systems.
Problem

Research questions and friction points this paper is trying to address.

Long-Horizon Multi-Agent Systems
Adversarial Stress-Testing
Persistent Deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

adversarial stress testing
long-horizon systems
multi-agent environments
resilience engineering
🔎 Similar Papers
No similar papers found.
D
Deepak Akkil
Emergence AI
T
Tamer Abuelsaad
Emergence AI
K
Karthik Vikram
Emergence AI
M
Matthew Pace
Emergence AI
Aditya Vempaty
Aditya Vempaty
Research Scientist, Emergence AI
Human-Machine Inference NetworksHuman Machine SystemsAI AgentsDecision MakingNetwork
S
Saahir Beotra
Emergence AI
Ravi Kokku
Ravi Kokku
Merlyn Mind Inc.
S
Satya Nitta
Emergence AI