Testing and Evaluation of Agentic AI Systems In Military Command and Control

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究针对军事指挥控制中代理AI系统的测试与评估问题,通过分析240个实践案例,提出保障声明并探讨现有方法的有效性。
📝 Abstract
Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discharged depends on their supporting assurance case, which requires three elements: claims specifying the conditions for acceptability, evidence bearing on those claims, and an argument connecting the two. Through a structured review of 240 documented Testing and Evaluation (T&E) practices, spanning eight evaluation dimensions and three lifecycle stages, we identify eight assumptions that established methods make about their test article, grouped into four clusters: system specifiability, stability, composability, and supervisability. Agentic properties weaken all eight assumptions. This erosion affects the argument connecting evidence to claims, not the claims or evidence themselves. As a result, test results may satisfy process requirements, but they do not warrant the inference from tested to fielded behavior. We derive ten assurance claims for the first three assumption clusters and assess whether current and emerging methods can address each, mapping operational consequences through five C2 scenarios. Supervisability is identified but not assessed here, since evidencing it depends on system stability results and human factors T&E methods beyond the present scope. The documented record does not support broad claims about system-level behavior, but narrower claims remain recoverable in principle, contingent on mature methods: bounded mission envelopes, trajectory-grounded correctness, executable runtime constraints, and characterized run-to-run variance. Part of the evidentiary burden shifts into deployment, making the determination to field a continuing act. Where evidence cannot be generated, the residual uncertainty can be governed through defined expiry conditions and assigned ownership.
Problem

Research questions and friction points this paper is trying to address.

Agentic AI
Military Command and Control
Testing and Evaluation
Assurance Case
System Stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic AI Systems
Assurance Claims
Testing and Evaluation (T&E) Practices
Supervisability
Deployment Evidence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
U
Ulysse Richard
Arcadia Impact, AI Governance Taskforce
H
Heather Frase
Veraitech
S
Sarah Cao
Arcadia Impact, AI Governance Taskforce; University of Oxford
D
Di Cooke
Arcadia Impact, AI Governance Taskforce; King’s College London
S
Sebastian Kwon
Arcadia Impact, AI Governance Taskforce; Atlantic Council
A
Adrianna Tan
Arcadia Impact, AI Governance Taskforce; Future Ethics Lab