🤖 AI Summary
This study addresses the insufficient research on risks to human agency and control arising from agent cognitive expansion. We propose a novel three-tiered cognitive risk analysis framework that systematically links physical, social, and self-referential cognition to threats against human autonomy. Through hierarchical cognitive modeling and systematic evaluation, this work elucidates the underlying risk mechanisms at each level and formulates targeted mitigation strategies. By bridging the critical gap in correlating cognitive capabilities with control risks, this research provides both theoretical foundations and practical pathways for enhancing agent controllability and ensuring the long-term safety of autonomous systems.
📝 Abstract
Frontier agentic systems powered by large language models (LLMs) exhibit human-like patterns of cognition. As these systems become deeply integrated across different domains, their cognitive engagement raises critical concerns for human society that remain insufficiently studied. To address this gap, we systematically analyze risks induced by expanding cognitive capabilities, following a three-level framework defined by their cognitive scope, from physical cognition to social cognition, and finally to self-referential cognition. We study their potential risks to human agency, autonomy, and control capability, corresponding to each cognitive level. We finally propose strategies to mitigate these risks and enhance the controllability of agentic AI systems, ensuring their long-term safe development.