A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Baselines

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过经典时间序列方法分析时空预测基准数据集,揭示了为何非空间线性模型表现优异,并建议减少对这些数据集的依赖,提倡更严格的统计评估。
📝 Abstract
Graph neural networks (GNNs) are routinely employed for short-range forecasting on multivariate time series with a spatial graph structure. Despite the availability of many alternative datasets, method innovations within this domain are predominantly assessed against a rather limited set of benchmark datasets, most notably Chickenpox, PedalMe, WikiMaths, METR-LA, and PEMS-BAY. The evaluation protocols contain baselines spanning from historical averages to classical machine learning approaches. These baselines often show competitive performance compared to GNNs. In the present work, we take a step back and analyse the benchmark datasets via classical time series methods to uncover why spatially-unaware linear models pose a stronger competitor than previously reported, casting further doubt on the discriminative reliability of the aforementioned widely adopted datasets. Our statistical analysis provides a toolset for identifying significant spatial and temporal correlations, while revealing a structural bias introduced by first-order differenced datasets. We therefore recommend reducing the over-reliance on such datasets for method comparison, and instead advocate for more rigorous statistical evaluation. By applying the results of our analysis to a simple hybrid model, we show how our methodology can lead to novel ways of developing GNN models
Problem

Research questions and friction points this paper is trying to address.

Spatiotemporal Forecasting
Benchmark Datasets
Graph Neural Networks
Linear Models
Statistical Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

spatiotemporal forecasting
graph neural networks
benchmark datasets
statistical analysis
hybrid model
🔎 Similar Papers
No similar papers found.