Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents
为解决动态未知环境中自主智能体的学习和对齐问题,研究提出通过内在动机引导探索学习,并借鉴儿童社会规范学习过程,构建动态教育环境以逐步实现智能体与人类目标的对齐。
为解决动态未知环境中自主智能体的学习和对齐问题,研究提出通过内在动机引导探索学习,并借鉴儿童社会规范学习过程,构建动态教育环境以逐步实现智能体与人类目标的对齐。
本文提出DARKSIDE方法,通过构建显式排除路径和担保轴来审计大型语言模型生成内容的连贯性,以解决模型对错误输入进行固化的问题。
This study investigates the emergence of joint agency and collective knowledge in multi-agent systems—specifically, how locally interacting agents generate group-level cognition and coordinated responses that transcend individual perceptual capacities. Method: Grounded in the active inference framework, we develop a flocking model integrating free-energy minimization with synergistic information decomposition to quantify inter-agent informational coupling dynamics. Contribution/Results: We demonstrate that strong informational coupling spontaneously induces statistical boundaries, enabling the group to function as an emergent agent with distinct perception, action, and internal states. Empirical results show that the collective exhibits faster, more coordinated responses to simulated predators, and encodes threat location information significantly exceeding the perceptual range of any single agent. This constitutes the first empirical validation—within the active inference paradigm—of implicitly encoded collective knowledge and collective sensitivity.
Existing cognitive models lack a unified computational framework to explain the dynamic interplay between semantic and episodic memory—and their synergistic roles in learning, recall, and imagination. To address this, we propose GENESIS: a neurobiologically grounded, dual-system generative architecture wherein semantic memory is modeled by a cortical variational autoencoder (VAE) for compression and generalization, while episodic memory is instantiated by a hippocampal retrieval-augmented generation (RAG) module supporting contextual binding and replay. GENESIS formalizes memory as constructive, active, and resource-constrained. It successfully reproduces core empirical phenomena—including semantic generalization, sequential recall, episodic distortion, and creative simulation—and quantifies how capacity limits degrade memory fidelity. By unifying semantic and episodic memory within a single generative framework, GENESIS advances a computationally precise, biologically plausible theory of declarative memory.
In predicting upper-limb lymphedema following breast cancer radiotherapy, existing interpretable AI models lack clinically intelligible quantitative explanations. This paper proposes a novel method integrating information retrieval (IR) metrics with rule-based modeling: it is the first to apply standard IR evaluation measures—such as Average Precision (AP) and Normalized Discounted Cumulative Gain (NDCG)—to quantify the contribution strength of individual clinical risk factors within rule-based predictions; further, it embeds an attribution-based explanation framework to enable comparable and verifiable assessment of feature importance. Experiments demonstrate substantial improvements in explanation intuitiveness and clinical consistency; a user study confirms that clinicians find the output significantly more comprehensible and actionable for decision support than conventional XAI methods. The core innovation lies in transferring IR’s ranking-evaluation paradigm to medical interpretability modeling, thereby bridging algorithmic interpretability and clinical utility.
为解决动态未知环境中自主智能体的学习和对齐问题,研究提出通过内在动机引导探索学习,并借鉴儿童社会规范学习过程,构建动态教育环境以逐步实现智能体与人类目标的对齐。
本文提出DARKSIDE方法,通过构建显式排除路径和担保轴来审计大型语言模型生成内容的连贯性,以解决模型对错误输入进行固化的问题。
This study investigates the emergence of joint agency and collective knowledge in multi-agent systems—specifically, how locally interacting agents generate group-level cognition and coordinated responses that transcend individual perceptual capacities. Method: Grounded in the active inference framework, we develop a flocking model integrating free-energy minimization with synergistic information decomposition to quantify inter-agent informational coupling dynamics. Contribution/Results: We demonstrate that strong informational coupling spontaneously induces statistical boundaries, enabling the group to function as an emergent agent with distinct perception, action, and internal states. Empirical results show that the collective exhibits faster, more coordinated responses to simulated predators, and encodes threat location information significantly exceeding the perceptual range of any single agent. This constitutes the first empirical validation—within the active inference paradigm—of implicitly encoded collective knowledge and collective sensitivity.
Existing cognitive models lack a unified computational framework to explain the dynamic interplay between semantic and episodic memory—and their synergistic roles in learning, recall, and imagination. To address this, we propose GENESIS: a neurobiologically grounded, dual-system generative architecture wherein semantic memory is modeled by a cortical variational autoencoder (VAE) for compression and generalization, while episodic memory is instantiated by a hippocampal retrieval-augmented generation (RAG) module supporting contextual binding and replay. GENESIS formalizes memory as constructive, active, and resource-constrained. It successfully reproduces core empirical phenomena—including semantic generalization, sequential recall, episodic distortion, and creative simulation—and quantifies how capacity limits degrade memory fidelity. By unifying semantic and episodic memory within a single generative framework, GENESIS advances a computationally precise, biologically plausible theory of declarative memory.
In predicting upper-limb lymphedema following breast cancer radiotherapy, existing interpretable AI models lack clinically intelligible quantitative explanations. This paper proposes a novel method integrating information retrieval (IR) metrics with rule-based modeling: it is the first to apply standard IR evaluation measures—such as Average Precision (AP) and Normalized Discounted Cumulative Gain (NDCG)—to quantify the contribution strength of individual clinical risk factors within rule-based predictions; further, it embeds an attribution-based explanation framework to enable comparable and verifiable assessment of feature importance. Experiments demonstrate substantial improvements in explanation intuitiveness and clinical consistency; a user study confirms that clinicians find the output significantly more comprehensible and actionable for decision support than conventional XAI methods. The core innovation lies in transferring IR’s ranking-evaluation paradigm to medical interpretability modeling, thereby bridging algorithmic interpretability and clinical utility.