FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决安全关键飞行预测中现有评估协议的不足,提出FLY-EVAL++,结合确定性验证与物理可行性、安全性约束进行多维度评分。
📝 Abstract
Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics, because predictions that are numerically close to the ground truth can still violate operational constraints, combine fields in physically inconsistent ways, or fail to produce usable structured outputs. Existing evaluation protocols do not measure these failure modes reliably. We propose FLY-EVAL++, an evidence-driven evaluation protocol that combines deterministic verification of protocol compliance, physical feasibility, and safety constraints with fixed rubric-guided aggregation into interpretable multi-dimensional scores. We instantiate FLY-EVAL++ for Flight Trajectory and Attitude Prediction (FTAP) by extending the PilotBench setting with history-conditioned and multi-step prediction tasks. Across 66 LLMs, safety compliance is the most discriminative dimension of model behavior: models with comparable predictive performance differ by more than 28 points in safety score, and we observe recurrent failures including safety violations under physically plausible predictions and instability in multi-step rollouts. These results show that evaluation in safety-critical domains should measure constraint satisfaction and structured validity explicitly rather than rely on accuracy-centric reporting alone.
Problem

Research questions and friction points this paper is trying to address.

large language models
safety-critical environments
evaluation protocol
operational constraints
physical feasibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

evidence-driven evaluation
safety constraints
physical feasibility
multi-dimensional scores
constraint satisfaction
💼 Related Jobs
No related jobs found.
Y
Yalun Wu
NExT++ Lab, School of Computing, National University of Singapore
Junfeng Fang
Junfeng Fang
National University of Singapore
Model EditingAI SafetyLLM ExplainabilityAI4Science
J
Jiawei Wang
NExT++ Lab, School of Computing, National University of Singapore
H
Haotian Liu
Xiamen University
Q
Qijun Yang
University of Manchester
Minghan Yang
Minghan Yang
Minghong Investment, Shanghai
optimizationmachine learning
Hongcheng Guo
Hongcheng Guo
School of Data Science, Fudan University
LLMsMultimodal LLMs
Zhoujun Li
Zhoujun Li
Beihang University
Artificial IntelligentNatural Language ProcessingNetwork Security
B
Boyang Wang
Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd., Beijing, China