🤖 AI Summary
This study challenges the conventional paradigm in time series forecasting that assumes longer historical context invariably improves performance, demonstrating instead that distant observations often introduce noise and induce inverse scaling. To address this issue, the authors propose RAFT (Retrieval-Augmented Forecasting Transformer), a framework that integrates a selective retrieval mechanism into foundation models to dynamically incorporate the most relevant historical segments as exogenous variables, thereby endowing the model with a context-aware inductive bias. Evaluated on continuous-context architectures such as PatchTST, RAFT achieves an MSE of 0.379 on ETTh1—significantly outperforming both long-context baselines and zero-shot models like Chronos and Moirai—while maintaining lower computational overhead. Notably, extending the input context to 3,000 steps increases prediction error by over 68%, underscoring the efficacy of selective context utilization.
📝 Abstract
Time Series Foundation Models (TSFMs) have borrowed the long context paradigm from natural language processing under the premise that feeding more history into the model improves forecast quality. But in stochastic domains, distant history is often just high-frequency noise, not signal. Hence, the proposed work tests whether this premise actually holds by running continuous context architectures (PatchTST included) through the ETTh1 benchmark. The obtained results contradict the premise: an inverse scaling law shows up clearly, with forecasting error rising as context gets longer. A 3,000-step window causes performance to drop by over 68%, evidence that attention mechanisms are poor at ignoring irrelevant historical volatility. Retrieval-Augmented Forecasting (RAFT) is evaluated as an alternative. RAFT achieves a mean squared error (MSE) of 0.379 with a fixed 720-step window and selective retrieval, outperforming both long-context configurations and zero-shot foundation models (Chronos, Moirai) despite requiring far less computation. In addition, the retrieval step injects only the most relevant historical segments as dynamic exogenous variables, which gives the model a context-informed inductive bias it cannot build on its own from raw sequences. Therefore, foundation models going forward need to shift architecturally toward selective retrieval.