🤖 AI Summary
Deep learning models are often assumed superior to classical statistical methods in terrestrial water storage (TWS) forecasting, yet their practical advantages remain inadequately validated under complex hydrological regimes driven jointly by natural variability and anthropogenic interventions.
Method: This study conducts a systematic, multi-scenario evaluation of LSTM and Temporal Fusion Transformer (TFT) against linear regression using the global HydroGlobe dataset, with rigorous out-of-sample validation across diverse hydroclimatic settings.
Contribution/Results: Linear regression consistently outperforms both deep learning models across most metrics and scenarios, demonstrating superior robustness and generalizability. These findings challenge the prevailing assumption that deep learning inherently yields better hydrological forecasts, underscoring the necessity of establishing strong, interpretable baselines—particularly linear models—for fair model assessment. The study further advocates developing a standardized global TWS benchmark dataset explicitly encoding coupled natural–human drivers to enable scientifically grounded, reproducible model evaluation.
📝 Abstract
Recent advances in machine learning such as Long Short-Term Memory (LSTM) models and Transformers have been widely adopted in hydrological applications, demonstrating impressive performance amongst deep learning models and outperforming physical models in various tasks. However, their superiority in predicting land surface states such as terrestrial water storage (TWS) that are dominated by many factors such as natural variability and human driven modifications remains unclear. Here, using the open-access, globally representative HydroGlobe dataset - comprising a baseline version derived solely from a land surface model simulation and an advanced version incorporating multi-source remote sensing data assimilation - we show that linear regression is a robust benchmark, outperforming the more complex LSTM and Temporal Fusion Transformer for TWS prediction. Our findings highlight the importance of including traditional statistical models as benchmarks when developing and evaluating deep learning models. Additionally, we emphasize the critical need to establish globally representative benchmark datasets that capture the combined impact of natural variability and human interventions.