🤖 AI Summary
This paper addresses the conceptual ambiguity and disciplinary fragmentation surrounding “AI safety.” It critically examines two prevailing tendencies in current research: (1) narrowing AI safety exclusively to speculative, long-term catastrophic risks, and (2) uncritically mapping it onto traditional safety engineering paradigms. Drawing on conceptual engineering methodology—integrating philosophical analysis, terminological critique, and cross-disciplinary empirical investigation—the authors propose a unified, normatively grounded definition centered on “preventing or mitigating harms caused by AI systems.” This definition explicitly encompasses empirically documented harms—including algorithmic bias, disinformation, and privacy violations—thereby ensuring inclusivity, historical continuity with prior safety scholarship, and actionable guidance for governance. For the first time, it achieves conceptual integration of AI safety across both descriptive and normative dimensions, advancing a consensus paradigm that evaluates AI governance through the lens of concrete, realized harms. (149 words)
📝 Abstract
The field of AI safety seeks to prevent or reduce the harms caused by AI systems. A simple and appealing account of what is distinctive of AI safety as a field holds that this feature is constitutive: a research project falls within the purview of AI safety just in case it aims to prevent or reduce the harms caused by AI systems. Call this appealingly simple account The Safety Conception of AI safety. Despite its simplicity and appeal, we argue that The Safety Conception is in tension with at least two trends in the ways AI safety researchers and organizations think and talk about AI safety: first, a tendency to characterize the goal of AI safety research in terms of catastrophic risks from future systems; second, the increasingly popular idea that AI safety can be thought of as a branch of safety engineering. Adopting the methodology of conceptual engineering, we argue that these trends are unfortunate: when we consider what concept of AI safety it would be best to have, there are compelling reasons to think that The Safety Conception is the answer. Descriptively, The Safety Conception allows us to see how work on topics that have historically been treated as central to the field of AI safety is continuous with work on topics that have historically been treated as more marginal, like bias, misinformation, and privacy. Normatively, taking The Safety Conception seriously means approaching all efforts to prevent or mitigate harms from AI systems based on their merits rather than drawing arbitrary distinctions between them.