🤖 AI Summary
This study addresses the challenge of quantifying uncertainty arising from internal climate variability, which is hindered by the prohibitive computational cost of running high-resolution, large-ensemble climate simulations. For the first time, a conditional variational autoencoder (cVAE) is applied to CMIP6 CanESM5 monthly-scale data, enhanced with an output noise injection mechanism to better capture multiscale climate variability. The proposed method generates physically consistent synthetic ensembles of arbitrary size from limited samples, accurately reproducing realistic teleconnection patterns and both low- and high-order statistical features—including extremes—even under climate conditions not present in the training data. This approach significantly improves the reliability of uncertainty assessments for both historical and future climate scenarios, while maintaining computational efficiency and mathematical interpretability.
📝 Abstract
Accurately quantifying uncertainty in predictions and projections arising from irreducible internal climate variability is critical for informed decision making. Such uncertainty is typically assessed using ensembles produced with physics based climate models. However, computational constraints impose a trade off between generating the large ensembles required for robust uncertainty estimation and increasing model resolution to better capture fine scale dynamics. Generative machine learning offers a promising pathway to alleviate these constraints. We develop a conditional Variational Autoencoder (cVAE) trained on a limited sample of climate simulations to generate arbitrary large ensembles. The approach is applied to output from monthly CMIP6 historical and future scenario experiments produced with the Canadian Centre for Climate Modelling and Analysis'(CCCma's) Earth system model CanESM5. We show that the cVAE model learns the underlying distribution of the data and generates physically consistent samples that reproduce realistic low and high moment statistics, including extremes. Compared with more sophisticated generative architectures, cVAEs offer a mathematically transparent, interpretable, and computationally efficient framework. Their simplicity lead to some limitations, such as overly smooth outputs, spectral bias, and underdispersion, that we discuss along with strategies to mitigate them. Specifically, we show that incorporating output noise improves the representation of climate relevant multiscale variability, and we propose a simple method to achieve this. Finally, we show that cVAE-enhanced ensembles capture realistic global teleconnection patterns, even under climate conditions absent from the training data.