Score
Performs probabilistic reliability analysis, producing outage probability estimates, reliability models, and uncertainty quantification for system availability.
To address conservative reliability assessment for safety-critical software under prior uncertainty, this paper proposes a robust Bayesian framework that computes the worst-case posterior predictive probability of fault-free operation—thereby yielding a conservative estimate of future reliability. Methodologically, software failures are modeled as a Bernoulli process, and the approach integrates set-based Bayesian inference with asymptotic analysis. Key contributions include: (1) the first closed-form analytical solution for the worst-case posterior predictive probability; (2) characterization of its asymptotic convergence properties; and (3) an extension of robust Bayesian theory, providing a rigorous mathematical foundation for quantifying worst-case behavior under prior uncertainty. The framework balances theoretical rigor with practical applicability, enabling high-assurance software reliability certification.
This study addresses the limitations of traditional Markovian approaches in availability analysis of repairable systems, which rely on the restrictive assumption of exponential distributions and thus fail to accurately capture real-world failure and repair time characteristics. To overcome this constraint, the work introduces the Lindley distribution—represented via phase-type approximation—into availability modeling for the first time, establishing a general analytical framework applicable to both single-component and n-component series-parallel systems. Closed-form expressions for time-dependent and steady-state availability are derived, along with an exact computation of mean time to repair. Numerical experiments demonstrate that incorporating non-exponential repair times significantly influences system reliability metrics, thereby underscoring the practical relevance and theoretical contribution of the proposed methodology.
Estimating extremely low-probability system failure events—particularly those arising from the tail regions of parameter distributions—remains a computationally prohibitive challenge. To address this, we propose the Tail-Stratified Sampling (TSS) estimator. TSS is the first method to enable adaptive, hierarchical, and direct sampling from parameter distribution tails, integrating stratified sampling theory, dynamic threshold partitioning, tail-oriented sampling, and multi-scale distribution reconstruction. Across analytical and high-dimensional numerical benchmarks, TSS achieves high-accuracy failure probability estimates using one to two orders of magnitude fewer system evaluations than conventional methods. Moreover, its estimator variance is reduced by one to two orders of magnitude, demonstrating superior computational efficiency and statistical robustness. The approach establishes a scalable, theoretically grounded paradigm for infrastructure reliability analysis under rare-event regimes.
Existing fault tree analysis (FTA) methods lack support for structured querying and precise modeling of uncertainty. This paper proposes stochastic fault tree logic (sfPFL), the first fault tree logic framework endowed with rigorous probabilistic semantics. By integrating a domain-specific language (DSL), sfPFL enables highly readable and expressive property specifications, and establishes a complete model-checking theory based on binary decision diagrams (BDDs). The approach supports natural, precise modeling, querying, and risk assessment of complex failure scenarios involving uncertainty. Experimental evaluation on COVID-19 transmission chains and oil–gas pipeline fault trees demonstrates that sfPFL significantly improves query efficiency for risk scenarios and enhances specification expressiveness, thereby enabling automated regulatory compliance verification.
Earth-entry capsules in NASA’s Mars Sample Return mission face a high risk of sample seal failure under extreme aerothermal and mechanical loads, jeopardizing the mission’s six-nines (0.999999) reliability target. Method: This study proposes a Bayesian Gaussian Process (BGP)-based probabilistic reliability assessment framework that tightly integrates uncertainty quantification, surrogate modeling, and probabilistic risk analysis. Contribution/Results: The resulting auditable and incrementally updatable statistical framework enables high-confidence verification of the six-nines reliability requirement. By jointly leveraging sparse physical test data and high-dimensional simulation outputs, the method significantly improves failure probability estimation accuracy for deep-space critical systems. Quantitative analysis confirms that the Earth-entry system meets mission-level reliability requirements. The approach provides NASA with a rigorous, transparent, and traceable decision-support foundation for the Mars Sample Return campaign.
This study addresses the critical need for reliable probabilistic forecasting of power network failures that accounts for both extreme and non-extreme weather events while quantifying uncertainty. The authors propose a unified probabilistic framework: non-extreme failures are modeled using multi-additive quantile regression with linear interpolation, while extreme events are captured via a discrete generalized Pareto distribution. By directly integrating ensemble numerical weather predictions, the method delivers up to four-day-ahead failure probability forecasts, further enhanced through probabilistic calibration to improve reliability. Evaluated on historical data from two UK distribution networks, the approach significantly outperforms existing methods. A real-world operational trial with Scottish Power Energy Networks from October 2024 to March 2025 demonstrates its practical utility in supporting maintenance decisions and reducing outage durations.
This work addresses the challenge of scarce fault data in safety-critical applications such as helicopter transmission systems by proposing an interpretable anomaly detection method that relies solely on healthy operational data. The approach employs Bayesian probabilistic modeling to learn the distribution of normal system states and introduces a tailored anomaly metric for real-time fault warning. By integrating uncertainty quantification with a visualization-based explanation mechanism, the method enhances the trustworthiness of diagnostic decisions. Experimental evaluation on both public predictive maintenance benchmarks and multi-year real-world helicopter transmission datasets demonstrates that the proposed technique achieves state-of-the-art detection performance while offering clear interpretability, making it well-suited for industrial settings demanding high reliability.
This study addresses the challenge of estimating time-varying outage risk and quantifying uncertainty in restoration processes using county-level power outage data. The authors propose a hierarchical Bayesian model that employs cubic B-spline basis functions to capture smooth, nonlinear restoration trajectories and adopts a Beta-Binomial likelihood to account for overdispersion in customer counts. Information sharing across geographic groups is achieved through shared hyperpriors. This approach provides, for the first time in outage restoration analysis, well-calibrated uncertainty quantification, enabling robust inference even under data sparsity. Experiments on diverse outage events in southern Wisconsin demonstrate that the model’s posterior mean reproduces trapezoidal AUC point estimates while yielding reliable uncertainty intervals—unattainable with deterministic methods—thereby offering actionable insights for emergency planning and decision-making.
Structural reliability analysis heavily relies on specialized expertise, which limits its broader engineering application. This work proposes a multi-agent large language model framework that, for the first time, integrates a fine-tuned Method Planner with a multi-agent architecture to automate the entire component-level reliability analysis pipeline—from natural language problem descriptions through modeling, method planning, code generation, execution, and result interpretation—while incorporating human verification at critical decision points. By delegating computations to validated deterministic solvers rather than relying on the LLM to generate numerical results directly, the system significantly enhances reproducibility and mitigates hallucination. Experimental results demonstrate that the proposed approach lowers the expertise barrier while preserving the accuracy and trustworthiness of the computational outcomes.
This study addresses the challenge of fault prediction in optical amplifiers by proposing a lightweight Transformer-based edge intelligence approach that leverages operational monitoring data to achieve high-accuracy remaining useful life estimation. The proposed method introduces, for the first time, a lightweight Transformer architecture into optical network operations and maintenance, enabling low-latency and efficient predictive maintenance. Experimental results demonstrate that the approach significantly enhances the availability and reliability of optical networks, while also validating the practical feasibility and effectiveness of deploying AI-driven models in autonomous optical networks.