FaCTR: Factorized Channel-Temporal Representation Transformers for Efficient Time Series Forecasting

📅 2025-06-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address architectural mismatch and over-parameterization in Transformers for time-series forecasting—caused by low input information density and complex cross-channel coupling—this paper proposes the lightweight Factorized Channel-Temporal Transformer (FACT). FACT models dynamic, symmetric cross-channel interactions via low-rank decomposition, and integrates temporal patch embedding, covariate encoding, and learnable gated fusion to jointly model static and dynamic multivariate dependencies. It introduces the first explicitly structured design for interpretable channel-wise influence analysis and self-supervised pretraining. Evaluated on 11 public benchmarks, FACT achieves state-of-the-art performance with a maximum parameter count of only ~400K—on average 50× smaller than comparable spatiotemporal Transformers—while significantly improving prediction accuracy, inference efficiency, and decision interpretability.

Technology Category

Application Category

📝 Abstract
While Transformers excel in language and vision-where inputs are semantically rich and exhibit univariate dependency structures-their architectural complexity leads to diminishing returns in time series forecasting. Time series data is characterized by low per-timestep information density and complex dependencies across channels and covariates, requiring conditioning on structured variable interactions. To address this mismatch and overparameterization, we propose FaCTR, a lightweight spatiotemporal Transformer with an explicitly structural design. FaCTR injects dynamic, symmetric cross-channel interactions-modeled via a low-rank Factorization Machine into temporally contextualized patch embeddings through a learnable gating mechanism. It further encodes static and dynamic covariates for multivariate conditioning. Despite its compact design, FaCTR achieves state-of-the-art performance on eleven public forecasting benchmarks spanning both short-term and long-term horizons, with its largest variant using close to only 400K parameters-on average 50x smaller than competitive spatiotemporal transformer baselines. In addition, its structured design enables interpretability through cross-channel influence scores-an essential requirement for real-world decision-making. Finally, FaCTR supports self-supervised pretraining, positioning it as a compact yet versatile foundation for downstream time series tasks.
Problem

Research questions and friction points this paper is trying to address.

Addresses inefficiency of Transformers in time series forecasting
Handles low information density and complex dependencies in time series
Proposes lightweight spatiotemporal Transformer with interpretable design
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lightweight spatiotemporal Transformer with structural design
Low-rank Factorization Machine for cross-channel interactions
Self-supervised pretraining for versatile downstream tasks
💼 Related Jobs
No related jobs found.
Y
Yash Vijay
C3 AI
H
Harini Subramanyan
C3 AI