Modular Foundation Models for Time-Series Perception in Digital Twins

📅 2026-07-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of task-specific, data-hungry, and poorly generalizable temporal-aware modules in digital twin and Prognostics and Health Management (PHM) systems, which also suffer from integration challenges. To overcome these issues, the authors propose a modular foundation model based on an ensemble of pretrained encoders. The model leverages self-supervised learning to acquire transferable temporal representations, employs a gating mechanism for dynamic encoder selection, and utilizes Transformer-based self-attention to enable cross-encoder interaction and representation fusion. Innovatively, it adopts a shared latent space alignment combined with an adaptive aggregation strategy, allowing lightweight multi-task adaptation and conditional computation while keeping the pretrained encoders frozen. The approach demonstrates superior performance on the ETT benchmark and validates its practical utility in an industrial virtual sensing application for hydro-generator rotor temperature monitoring.
📝 Abstract
Engineering Digital Twins and Prognostics and Health Management (PHM) systems rely on robust perception modules to extract actionable information from heterogeneous and non-stationary time-series data. However, most existing approaches remain task-specific, data-hungry, and difficult to integrate into scalable monitoring and decision-making pipelines. Moreover, purely data-driven models often lack robustness and transferability across varying operating conditions. To address these challenges, this paper proposes a modular foundation model for time-series perception based on a collection of pretrained representation encoders. The framework leverages self-supervised learning on heterogeneous datasets to learn transferable and task-agnostic representations, which can be reused across multiple PHM tasks. A gating mechanism is introduced to dynamically select relevant encoders for a given target dataset, enabling conditional computation and adaptive model composition. The selected representations are projected into a shared latent space and aggregated using a Transformer-based self-attention module that explicitly models cross-encoder interactions. The resulting architecture supports multiple downstream tasks, including imputation, long-term forecasting, and few-shot learning, through lightweight task-specific heads, while keeping pretrained encoders frozen during adaptation. Extensive ablation studies demonstrate the complementary roles of self-supervised pretraining, encoder selection, representation alignment, and adaptive aggregation. Experimental results on the ETT benchmark show competitive performance across tasks, while a real-world industrial case study on virtual sensing for hydro-generator rotor temperature highlights the practical relevance of the approach.
Problem

Research questions and friction points this paper is trying to address.

time-series perception
digital twins
prognostics and health management
foundation models
transferability
Innovation

Methods, ideas, or system contributions that make the work stand out.

modular foundation model
self-supervised learning
time-series perception
adaptive encoder selection
cross-encoder attention
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Q
Quang Hung Pham
Hydro-Québec
R
Ryad Zemouri
Hydro-Québec
M
Martin Gagnon
Hydro-Québec
L
Luc Vouligny
Hydro-Québec