🤖 AI Summary
Current AI red-teaming practices overemphasize model-level vulnerabilities while neglecting emergent risks arising from interactions among models, users, and socio-technical environments—resulting in governance lagging behind the real-world deployment of generative AI. Method: We propose a “two-tiered red-teaming framework”: a macro-level tier spanning the AI lifecycle and integrating technical, organizational, and societal dimensions; and a micro-level tier preserving model-specific vulnerability detection. Our approach synthesizes systems theory, cybersecurity practice, and interdisciplinary collaboration to establish a dynamic, multi-layered risk identification and assessment system. Contribution/Results: This work is the first to institutionalize systems thinking in AI red-teaming, delivering an actionable implementation framework and practical guidelines. It shifts AI governance from isolated defect detection toward holistic, systemic risk mitigation—advancing proactive, context-aware safety assurance for deployed generative AI systems.
📝 Abstract
Red teaming has evolved from its origins in military applications to become a widely adopted methodology in cybersecurity and AI. In this paper, we take a critical look at the practice of AI red teaming. We argue that despite its current popularity in AI governance, there exists a significant gap between red teaming's original intent as a critical thinking exercise and its narrow focus on discovering model-level flaws in the context of generative AI. Current AI red teaming efforts focus predominantly on individual model vulnerabilities while overlooking the broader sociotechnical systems and emergent behaviors that arise from complex interactions between models, users, and environments. To address this deficiency, we propose a comprehensive framework operationalizing red teaming in AI systems at two levels: macro-level system red teaming spanning the entire AI development lifecycle, and micro-level model red teaming. Drawing on cybersecurity experience and systems theory, we further propose a set of recommendations. In these, we emphasize that effective AI red teaming requires multifunctional teams that examine emergent risks, systemic vulnerabilities, and the interplay between technical and social factors.