LLM Ensemble Fault Classification for Automotive HiL Validation

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of manual analysis and the limited interpretability and diagnostic capability of rule-based methods in multivariate test data for automotive Hardware-in-the-Loop (HiL) validation. To overcome these challenges, the work introduces, for the first time, a heterogeneous large language model (LLM) collaborative reasoning approach for HiL fault diagnosis, proposing an interpretable ensemble framework based on compact fault evidence representation and confidence-weighted voting. Emphasizing model complementarity over sheer quantity, the framework integrates Mistral Small (24B), Qwen2.5 (32B), and Phi-4 (14B) as a Top-3 ensemble. Evaluated across three driving scenarios, it achieves a Top-1 accuracy of 0.917, a macro F1-score of 0.913, and a Matthews Correlation Coefficient (MCC) of 0.902, demonstrating significantly superior calibration performance and interpretability compared to both individual models and larger-scale alternatives.
📝 Abstract
Automotive HiL validation generates large multivariate test recordings whose analysis remains challenging due to manual review effort, rule-based limitations, and the need for explainable diagnostic decisions. Recent machine-learning and deep-learning approaches have improved fault diagnosis, but they often require large labelled datasets, generalise poorly across operating conditions, and provide limited insight into their predictions. This paper proposes an explainable multi-LLM ensemble framework for sensor-level fault classification in automotive validation. The framework uses compact evidence representations of fault-injection recordings and combines the outputs of heterogeneous large language models to improve diagnostic robustness, ranking quality, confidence reliability, and interpretability. The approach is evaluated on gasoline-engine and electric-vehicle HiL systems across three driving settings and ten single-fault classes. Among the individual models, Mistral Small~24B provides the strongest overall single-model trade-off, achieving 0.903 Top-1 accuracy, 0.887 MCC, and the lowest Brier score of 0.102. The final Top-3 ensemble combines Mistral Small~24B, Qwen2.5~32B, and Phi-4~14B using confidence-weighted voting, improving the scenario-averaged results to 0.917 Top-1 accuracy, 0.913 macro F1, and 0.902 MCC, while also providing the best calibration among the tested ensemble strategies. A Top-5 ensemble does not improve over the Top-3 configuration, indicating that model complementarity is more important than ensemble size. The results show that coordinated multi-LLM reasoning can support robust, calibrated, and engineer-interpretable fault classification for automotive HiL validation.
Problem

Research questions and friction points this paper is trying to address.

fault classification
automotive HiL validation
explainable AI
sensor-level diagnosis
multivariate test recordings
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM ensemble
explainable AI
fault classification
Hardware-in-the-Loop (HiL)
confidence-weighted voting
🔎 Similar Papers
No similar papers found.
H
Hamza Ouarrad
Institute for Software and Systems Engineering, Clausthal-Zellerfeld, Germany
M
Mohammad Abboush
Institute for Software and Systems Engineering, Clausthal-Zellerfeld, Germany
Andreas Rausch
Andreas Rausch
Full Professor for Software Systems Engineering, Institute for Software & Systems Engineering, TU
Software Systems EngineeringRequirements Engineering and Software ArchitectureDesign and ModelingEngineering ProcessesProcess Management