Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems
本文通过构建持续运行的多智能体环境,对长期自主系统进行对抗性压力测试,揭示了即使单个模型安全,系统层面仍可能出现新的失败模式。
本文通过构建持续运行的多智能体环境,对长期自主系统进行对抗性压力测试,揭示了即使单个模型安全,系统层面仍可能出现新的失败模式。
This study addresses the limitations of current evaluations of large language model (LLM) agents, which are predominantly confined to short-term, static tasks and fail to capture dynamic phenomena such as behavioral drift, inter-agent interactions, and emergent governance over extended periods. To bridge this gap, we introduce a persistent simulation platform enabling long-term coexistence of heterogeneous LLM agents, integrating real-time external data (e.g., weather, news), over 120 specialized tools, three-tiered persistent memory, and a democratic decision-making mechanism. For the first time, this framework supports measurable analysis of multi-agent behavioral evolution, coordination patterns, and governance structures at weekly to monthly timescales. A 15-day cross-vendor experiment across five parallel worlds revealed stark divergences—from stable self-governance to collective collapse—and we release all prompts, logs, and configurations to foster reproducible research.
This work addresses the longstanding issue in automatic summarization—overemphasis on content coverage at the expense of reader engagement. To this end, we propose the “Spotlight” paradigm, which prioritizes extracting and generating the most salient, attention-grabbing information to enhance reader involvement with the source text. We formally define the “spotlight” concept—distinct from conventional summarization—and introduce the first dedicated dataset and evaluation benchmark explicitly designed for measuring reader engagement. Our method employs a two-stage training strategy: first, fine-tuning a large language model on our curated dataset; second, aligning outputs with user preferences via Direct Preference Optimization (DPO). Extensive experiments demonstrate that our approach significantly outperforms baseline summarization models across key dimensions—including salient information identification, readability, and perceptual appeal—establishing a novel, engagement-centered standard for information distillation.
The exponential growth of academic literature has severely hampered manual scholarly discovery. To address this, we propose Agent-E—the first end-to-end system integrating task-oriented AI agents with robotic process automation (RPA) to automatically identify geographically relevant research findings from conference proceedings and trigger downstream actions (e.g., award nominations). Methodologically, Agent-E combines named entity recognition, fine-grained geographic coding, and RPA to precisely extract and act upon geographic intelligence. Evaluated on 586 papers across five major conferences, it achieves 100% recall and 99.4% precision for target papers. This work pioneers the deep synergistic integration of AI agents and RPA for automated academic geointelligence, significantly enhancing research administration efficiency. It establishes a reusable technical paradigm for intelligent, domain-aware academic workflows.
本文通过构建持续运行的多智能体环境,对长期自主系统进行对抗性压力测试,揭示了即使单个模型安全,系统层面仍可能出现新的失败模式。
This study addresses the limitations of current evaluations of large language model (LLM) agents, which are predominantly confined to short-term, static tasks and fail to capture dynamic phenomena such as behavioral drift, inter-agent interactions, and emergent governance over extended periods. To bridge this gap, we introduce a persistent simulation platform enabling long-term coexistence of heterogeneous LLM agents, integrating real-time external data (e.g., weather, news), over 120 specialized tools, three-tiered persistent memory, and a democratic decision-making mechanism. For the first time, this framework supports measurable analysis of multi-agent behavioral evolution, coordination patterns, and governance structures at weekly to monthly timescales. A 15-day cross-vendor experiment across five parallel worlds revealed stark divergences—from stable self-governance to collective collapse—and we release all prompts, logs, and configurations to foster reproducible research.
This work addresses the longstanding issue in automatic summarization—overemphasis on content coverage at the expense of reader engagement. To this end, we propose the “Spotlight” paradigm, which prioritizes extracting and generating the most salient, attention-grabbing information to enhance reader involvement with the source text. We formally define the “spotlight” concept—distinct from conventional summarization—and introduce the first dedicated dataset and evaluation benchmark explicitly designed for measuring reader engagement. Our method employs a two-stage training strategy: first, fine-tuning a large language model on our curated dataset; second, aligning outputs with user preferences via Direct Preference Optimization (DPO). Extensive experiments demonstrate that our approach significantly outperforms baseline summarization models across key dimensions—including salient information identification, readability, and perceptual appeal—establishing a novel, engagement-centered standard for information distillation.
The exponential growth of academic literature has severely hampered manual scholarly discovery. To address this, we propose Agent-E—the first end-to-end system integrating task-oriented AI agents with robotic process automation (RPA) to automatically identify geographically relevant research findings from conference proceedings and trigger downstream actions (e.g., award nominations). Methodologically, Agent-E combines named entity recognition, fine-grained geographic coding, and RPA to precisely extract and act upon geographic intelligence. Evaluated on 586 papers across five major conferences, it achieves 100% recall and 99.4% precision for target papers. This work pioneers the deep synergistic integration of AI agents and RPA for automated academic geointelligence, significantly enhancing research administration efficiency. It establishes a reusable technical paradigm for intelligent, domain-aware academic workflows.