Agentic Microphysics: A Manifesto for Generative AI Safety
Current AI safety research struggles to explain how local interactions among agents endowed with planning, memory, and sustained interaction capabilities can lead to emergent macro-level unsafe behaviors. This work proposes an “agent microphysics” analytical framework that shifts the focus of safety analysis from individual agents or aggregate outcomes to interpretable and intervenable micro-level interaction dynamics under specific protocols. By integrating generative modeling, multi-agent protocol design, causal identification, and threshold detection, the framework enables the reconstruction of collective risk-generation mechanisms, identification of sufficient conditions and critical thresholds, and formulation of effective interventions. This approach establishes a novel paradigm for ensuring the safety of highly autonomous AI systems grounded in mechanistic understanding of their microscopic behaviors.