reliability assessment

Evaluating and validating system reliability by measuring, calibrating, and comparing model or system outputs against benchmarks and human judgments, and assessing operational properties (e.g., grid flexibility, recall, token costs, robustness) to inform planning and design.

reliabilityassessment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.98
Aug 01, 2026Aug 01, 2026
Career
Value
No comparison yet
$203K/year
Aug 01, 2026Aug 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

In the pre-prototype phase of complex novel systems, the absence of empirical data impedes rigorous assessment of simulation model credibility. Method: This paper proposes a physics-fidelity-based model trust evaluation method that bypasses reliance on real-world measurements. Instead, it quantifies model applicability by systematically analyzing the completeness of represented physical phenomena, the mathematical complexity of their formulation, and the fidelity of emergent behavior modeling. Contribution/Results: The approach enables objective, quantitative ranking of multiple candidate models under data-scarce conditions—thereby significantly enhancing the reliability of simulation-driven decisions during early-stage design. It establishes both theoretical foundations and practical tools for model-based design in high-uncertainty scenarios, advancing trustworthy digital twin development and physics-informed simulation validation.

Evaluating trustworthiness of simulation models for complex systemsSelecting appropriate physics-based models for design decisionsValidating models without real-world data in pre-prototype stages

This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.

calibrationdistributional validationprobabilistic forecasting

This paper addresses the reliability of calibration evaluation for machine learning models, identifying systematic biases in the widely used Expected Calibration Error (ECE) under distributional shift and varying binning strategies. Methodologically, it clarifies the logical hierarchy among multi-level calibration definitions, and systematically exposes ECE’s limitations through visualization, binning-based statistical analysis, and theoretical derivation—demonstrating its failure to satisfy key requirements of robustness and consistency in calibration assessment. Building on this critique, the paper introduces and explicates emerging calibration paradigms—including distribution-level and instance-level calibration—alongside their corresponding evaluation methodologies, thereby constructing a rigorous, interpretable, and practice-oriented calibration knowledge framework. The results equip researchers with principled guidance for selecting appropriate evaluation metrics and advance calibration assessment from ad hoc, heuristic practices toward standardization and formalization.

Evaluation MetricsLimitationsMachine Learning Calibration

This study addresses a critical limitation in existing simulation credibility assessment approaches, which predominantly focus on individual models and thus fail to capture the reliability of complex, multi-model architectures. Moving beyond the single-model evaluation paradigm, this work redefines trustworthiness at the architectural level and proposes a multidimensional framework that integrates sensitivity analysis, expert knowledge, explainable artificial intelligence, and complex network modeling. Through a systematic comparison of diverse methodologies across dimensions such as methodological rigor, generalizability, and computational resource demands, the research offers both theoretical foundations and practical guidance for constructing high-assurance simulation architectures.

assembly credibilitymodel credibilitysimulation architecture

Axioms for Model Fidelity Evaluation

Jul 30, 2025
ET
Evan Taylor
🏛️ Clemson University

Model fidelity—the degree of correspondence between simulation and reality—lacks a formal, axiomatic foundation in digital engineering, resulting in ambiguous evaluation criteria and poor cross-domain comparability. Method: This paper introduces the first rigorous, verifiable theoretical framework for fidelity assessment, grounded in seven foundational axioms encompassing consistency, measurability, scale invariance, and other essential properties; the framework enables formal verification and comparative analysis of fidelity metrics. Empirical validation is conducted via integration into ground-vehicle modeling, demonstrating feasibility and practical guidance within existing evaluation paradigms. Contribution/Results: The work fills a critical theoretical gap in fidelity science and establishes a universal, standards-ready paradigm for fidelity assessment—directly advancing digital twin development, simulation verification and validation (V&V), and model-based systems engineering. It further provides a clear, principled roadmap for future methodological evolution and standardization.

Addressing ambiguity in simulation-reality consistency assessmentDefining rigorous axioms for model fidelity evaluationEstablishing foundations for future fidelity frameworks

Latest Papers

What's happening recently
View more

This study addresses critical validity and reliability deficiencies in current AI agent evaluation benchmarks, which severely undermine deployment decisions and compliance judgments. Through benchmark auditing, LLM-based simulator calibration experiments, and psychometric analyses—including intraclass correlation (ICC) and Cronbach’s α—the work quantifies, for the first time, the multiplicative error effects across task generation, user simulation, and scoring judgment, yielding a total validity upper-bound formula: $V_{\text{total}} \leq V_1 \times V_2 \times V_3$. The analysis reveals that 82% of widely used benchmarks either omit or misuse rater reliability metrics. In response, the paper proposes eight psychometrically grounded evaluation guidelines, specifying concrete thresholds and mandatory reporting requirements to substantially enhance the credibility and scientific rigor of AI agent assessments.

Agentic AIbenchmarkingevaluation validity

Current agent leaderboards erroneously conflate task-specific performance with general capability, yielding unreliable deployment decisions. This work introduces four-facet generalizability theory to agent evaluation for the first time, employing variance decomposition to demonstrate that leaderboard rankings predominantly reflect task-specific expertise rather than general competence: main-effect variance is negligible, while interaction effects dominate. We estimate variance components using Henderson Method-I, REML (via lme4), and Bayesian binomial GLMMs, and analyze failure modes through MAST taxonomy–based classification. Experiments across three major benchmarks reveal that evaluation reliability collapses on difficult tasks and that training-unit reliability negatively correlates with holdout reliability. To address these issues, we propose a reliability assessment framework tailored for enterprise deployment and introduce the DDR reporting standard. All code and data are publicly released to enable cross-benchmark diagnostic transfer.

Agent EvaluationDeployment Decision ReliabilityGeneralizability Theory

Low-granularity operational data can lead to overly optimistic assessments of autonomous driving software reliability, thereby undermining the credibility of safety certification. This work proposes a systematic approach based on Conservative Bayesian Inference (CBI) to quantify, for the first time, the adverse impact of insufficient data fidelity on the robustness of reliability claims. By integrating statistical robustness analysis with software reliability modeling, the study demonstrates that even conservative inference strategies may yield misleading conclusions when applied to low-fidelity data. The paper establishes the first conservative estimation framework that explicitly accounts for the influence of data granularity on reliability assessment, highlighting the critical importance of high-fidelity operational data in safety certification of autonomous driving systems.

autonomous vehiclesfailure data granularityoperational-data fidelity

Current AI evaluation frameworks overemphasize output correctness while neglecting the resource costs required to verify errors in real-world deployment, allowing high accuracy metrics to mask substantial verification burdens. This work introduces verification-cost errors (VCEs)—errors that cannot be detected by a specified proportion of validators within a given verification budget—thereby shifting the paradigm from defining errors solely by output properties to centering on their detectability during verification. Through an operational definition, verification budget modeling, and user studies, we empirically demonstrate in code generation and multimodal document understanding tasks that high benchmark accuracy can coexist with significant verification effort, underscoring that correctness alone is insufficient to reflect system reliability in practical settings.

AI reliabilityevaluation metricsoutput correctness

Current computer-using agent (CUA) benchmarks rely on fragile scripted evaluators that frequently produce erroneous failure judgments, obscuring true performance bottlenecks. This work proposes the first reliability-focused evaluation framework encompassing the entire pipeline—from task construction and trajectory observation to scoring and reporting—and introduces a three-tier failure diagnosis taxonomy. Through manual auditing and attribution analysis of 150 publicly reported failure trajectories, we find that 15.3% of failure labels are incorrect, with 10.7% stemming from evaluator misjudgment and 4.7% arising from task design flaws. Building on these insights, we derive phased design principles for long-horizon CUA evaluation, substantially improving assessment accuracy and interpretability.

benchmarkingcomputer-use agentsevaluation reliability

Hot Scholars

CH

Chang Hee Lee

Associate Professor, KAIST
Interaction DesignDesign EngineeringHuman-Computer InteractionUser Experience
TK

Taeyong Kim

Department of Civil Systems Engineering, Ajou University
Structural ReliabilityDisaster ResilienceMachine LearningEarthquake Engineering
JO

Joe O'Brien

Institute for AI Policy and Strategy (IAPS)
artificial intelligenceAI auditingAI risk managementAI governance
BA

Basant Agarwal

Central University of Rajasthan
Deep LearningNatural Language ProcessingMachine LearningSentiment Analysis
JW

Jun Wang

Assistant Professor of Mechanical Engineering, Santa Clara University
Data-Driven Design and ManufacturingPhysics-Driven DesignDesign for Additive ManufacturingMetamaterials Design