EMTSF:Extraordinary Mixture of SOTA Models for Time Series Forecasting
Challenged by the limited efficacy of Transformers in time-series forecasting (TSF), the insufficient robustness of LLM-based approaches, and the over-dominance of recent observations, this paper proposes the first Transformer-gated Mixture-of-Experts (MoE) framework integrating multiple state-of-the-art paradigms. The framework unifies four heterogeneous models—xLSTM, an enhanced linear model, PatchTST, and minGRU—under a learnable Transformer-based gating network for dynamic expert weighting. It further introduces a recency-prioritized temporal weighting scheme to strengthen local dynamics modeling. Distinct from existing MoE methods, this work achieves cross-architectural complementarity within a single unified architecture, significantly improving both accuracy and robustness. Extensive experiments demonstrate consistent superiority over leading TSF models—including TimeLLM—across multiple standard benchmarks, empirically validating the effectiveness of heterogeneous model collaboration.