🤖 AI Summary
Cloud platforms supporting digital twin (DT) applications face reliability bottlenecks—including service outages, resource contention, and system fragility—while existing research lacks effective fault-tolerant and self-healing mechanisms for DTs. To address this, we propose a highly reliable DT cloud management model. First, we introduce a novel federated learning–based resource estimation mechanism incorporating cosine similarity to enable accurate cross-domain workload forecasting. Second, we design a root-cause diagnosis method leveraging frequent sequence pattern mining, coupled with a dynamic virtual machine (VM) self-healing scheduling strategy. Evaluated in multi-stakeholder manufacturing and processing scenarios, our approach improves service availability by 13.2% over baseline methods, significantly extends mean time between failures (MTBF), and markedly reduces mean time to repair (MTTR), thereby ensuring stable operation of mission-critical DT applications.
📝 Abstract
Digital twins (DTs), integral to cloud platforms, bridge physical and virtual worlds, fostering collaboration among stakeholders in manufacturing and processing. However, the cloud platforms face challenges such as service outages, vulnerabilities, and resource contention, hindering critical DT application development. The existing research works have limited focus on reliability and fault tolerance in DT processing. In this context, this article proposed a novel self-healing and fault-tolerant cloud-based digital twin processing management (SF-DTM) model. It employs collaborative DT tasks resource requirement estimation unit that utilizes newly devised federated learning with cosine similarity integration. Furthermore, SF-DTM incorporates a self-healing fault-tolerance strategy employing a frequent sequence fault-prone pattern analytics unit for deciding the most admissible virtual machine (VM) allocation. The implementation and evaluation of the SF-DTM model using real traces demonstrates its effectiveness and resilience, revealing improved availability, higher mean time between failure, and lower mean time to repair compared with non-SF-DTM approaches, enhancing collaborative DT application management. SF-DTM improved the services availability up to 13.2% over non-SF-DTM-based DT processing.