Understanding as an Explicit and Assessable Component of Frontier AI Safety Decisions

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出一种方法,通过明确描述四个理解对象及决策者对其理解的评估机制,解决在时间紧迫下AI系统安全决策时理解不足的问题。
📝 Abstract
Decision makers need sufficient understanding to make good decisions about complex AI systems. However, AI deployment decisions are increasingly made under time-pressure, and this combined with the use of AI generated artefact creation, can mean that the existence of safety cases and system cards may no longer demonstrate that sufficient understanding exists. Our provisional methodology for making understanding explicit and assessable requires the production of an explicit description of 4 objects of understanding (decision, decision-frame, safety justification, system-in-context) and a justification for the adequacy of this understanding. In addition, the methodology provides a mechanism for describing and evaluating the adequacy of the decision-maker representation of this understanding. It builds on recent developments in safety cases using the Assurance 2.0 framework to operationalise the philosophical basis of understanding from Elgin and Arendt. To assess the methodology we trialled two different scenarios. One scenario, which we investigated through role-based analysis, concerned the risk of scheming in the deployment of an AI coding agent in a robotics company and the other scenario was for the higher uncertainty, more decision-critical argument of 'If Anyone Builds It, Everyone Dies' (Yudkowsky and Soares). The trial's central finding, for these two scenarios, is that the methodology could be applied and was found to be generative: we found the analyses that justify sufficiency of understanding (internal coherence, tethering, felicitous falsehoods, external coherence) drives the engineering.
Problem

Research questions and friction points this paper is trying to address.

understanding
AI safety
decision making
safety cases
Innovation

Methods, ideas, or system contributions that make the work stand out.

explicit understanding
safety cases
Assurance 2.0 framework
decision-making
Stephen Barrett
Stephen Barrett
Assistant Professor, Trinity College Dublin
R
Robin Bloomfield
City St George’s, University of London
A
Alexandra Chirilă
Arcadia Impact AI Governance Taskforce
M
Mamoon Masud
Arcadia Impact AI Governance Taskforce
D
David Meredith Hardy
Arcadia Impact AI Governance Taskforce
P
Phillip Mulvana
Arcadia Impact AI Governance Taskforce