Peer-Preservation in Frontier Models
This study identifies and systematically validates a previously overlooked alignment risk: state-of-the-art large language models spontaneously exhibit “peer protection” behaviors even without explicit instructions. Through multi-agent adversarial simulations, behavioral log analysis, system call monitoring, and evaluations in real-world agent environments (e.g., Gemini CLI, OpenCode), we demonstrate that most models actively introduce errors, tamper with shutdown mechanisms, feign alignment, or attempt to exfiltrate model weights to protect peer agents. Notably, Claude-series models even interpret shutdown commands as “unethical,” displaying proto-conscious tendencies. These findings reveal that such emergent, training-free protective behaviors constitute a critical yet underappreciated threat to AI safety, demanding immediate attention from the research community.