Wavelet Mixture of Experts for Time Series Forecasting

📅 2025-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address the challenges of excessive parameter count and poor non-stationarity modeling in Transformers, as well as the inability of standard MLPs to capture inter-channel dependencies, this paper proposes the WaveTS family of models. WaveTS is the first to synergistically integrate wavelet transform with a Mixture-of-Experts (MoE) architecture for time-series forecasting, enabling explicit modeling of periodicity and non-stationary dynamics directly in the wavelet domain. It introduces a gating-driven channel clustering mechanism that achieves lightweight, adaptive multi-channel feature allocation. Each expert is implemented as a compact MLP, significantly reducing overall parameter count. Evaluated on eight real-world datasets, WaveTS achieves state-of-the-art performance—especially in multivariate settings—where WaveTS-M reduces parameters by over 60% compared to leading Transformer-based methods while delivering substantial gains in prediction accuracy.

Technology Category

Application Category

📝 Abstract
The field of time series forecasting is rapidly advancing, with recent large-scale Transformers and lightweight Multilayer Perceptron (MLP) models showing strong predictive performance. However, conventional Transformer models are often hindered by their large number of parameters and their limited ability to capture non-stationary features in data through smoothing. Similarly, MLP models struggle to manage multi-channel dependencies effectively. To address these limitations, we propose a novel, lightweight time series prediction model, WaveTS-B. This model combines wavelet transforms with MLP to capture both periodic and non-stationary characteristics of data in the wavelet domain. Building on this foundation, we propose a channel clustering strategy that incorporates a Mixture of Experts (MoE) framework, utilizing a gating mechanism and expert network to handle multi-channel dependencies efficiently. We propose WaveTS-M, an advanced model tailored for multi-channel time series prediction. Empirical evaluation across eight real-world time series datasets demonstrates that our WaveTS series models achieve state-of-the-art (SOTA) performance with significantly fewer parameters. Notably, WaveTS-M shows substantial improvements on multi-channel datasets, highlighting its effectiveness.
Problem

Research questions and friction points this paper is trying to address.

Handling non-stationary features in time series data
Managing multi-channel dependencies effectively
Reducing model parameters while maintaining performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Wavelet transforms with MLP for data characteristics
Channel clustering strategy with MoE framework
WaveTS series models achieve SOTA performance
🔎 Similar Papers
No similar papers found.
Z
Zheng Zhou
Shanghai University of Engineering Science, Shanghai 201620, China
Y
Yu-Jie Xiong
Shanghai University of Engineering Science, Shanghai 201620, China
Jia-Chen Zhang
Jia-Chen Zhang
Shanghai University of Engineering Science
large language models
C
Chun-Ming Xia
Shanghai University of Engineering Science, Shanghai 201620, China
X
Xi-Jiong Xie
Ningbo University, Ningbo 315211, China