Dual-Forecaster: A Multimodal Time Series Model Integrating Descriptive and Predictive Texts
Existing unimodal time series models rely solely on numerical data, suffering from semantic sparsity; while multimodal approaches incorporate textual information, they typically leverage only unidirectional text—either historical or future—and lack fine-grained modeling of text–time semantics, temporal dynamics, and causal relationships. To address these limitations, we propose a bidirectional text-driven forecasting paradigm that jointly integrates descriptive historical text and predictive future text for the first time. We design a three-stage cross-modal alignment module—encompassing semantic, temporal, and causal alignment—leveraging a large language model for text encoding and a dedicated time series feature extractor. Extensive experiments across 15 multivariate time series benchmarks demonstrate that our method consistently matches or surpasses state-of-the-art approaches, validating the substantial performance gains enabled by bidirectional textual integration.