TERRA: A Hierarchical Parallel Training and Memory Orchestration Framework for High-Resolution AI-based Earth Modeling

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses GPU memory bottlenecks and inadequate distributed system support in training high-resolution AI Earth prediction models by proposing a hierarchical parallel training framework. The approach employs a sampling window sequence-aware tensor parallelism strategy to preserve spatial continuity, integrated with rollout-aware checkpointing and budget-constrained activation offloading to optimize memory orchestration. This method effectively resolves computational and storage challenges in long-sequence fine-tuning. Evaluated on a 96-GPU H200 cluster, the framework supports training an 11.4-billion-parameter model, achieving a peak performance of 39.76 PFLOPS with strong and weak scaling efficiencies of 65% and 94.1%, respectively, while reducing peak GPU memory usage by over 32%.
📝 Abstract
Training high-resolution AI-based Earth forecasting models is memory-intensive. Window-based Swin Transformers reduce the quadratic cost of global attention, but existing distributed systems such as AERIS primarily target pixel-level models and do not jointly support convolutional sampling modules and shifted-window execution. Long-lead rollout finetuning further increases activation memory. To address these challenges, we present TERRA, a hierarchical parallel training framework for high-resolution Earth forecasting. TERRA introduces Sampling-Aware Window, Sequence, and Tensor Parallelism (SAWSTP), which preserves spatially contiguous layouts for sampling modules and routes tokens into topology-aware ragged window layouts for Transformer execution. For long-lead finetuning, Memory Orchestration (MO) provides rollout-aware checkpoint planning and combines input buffering with budget-constrained activation offloading. Experiments on the $1/12^\circ$ GLORYS-based Wenhai workload show that TERRA supports models with up to 11.4B parameters on 96 H200 GPUs and sustains up to $39.76$ PFLOPS, achieving $65.0\%$ strong-scaling and $94.1\%$ weak-scaling efficiency. Compared with checkpoint-only policies, MO further reduces peak allocated GPU memory by $32.2\%$--$51.8\%$ with at most $20.0\%$ step-time overhead, which makes finetuning with smaller patch sizes and longer rollouts feasible for improved forecasting accuracy.
Problem

Research questions and friction points this paper is trying to address.

High-Resolution Earth Modeling
Memory-Intensive Training
Distributed Systems
Long-lead Rollout Finetuning
Activation Memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sampling-Aware Window Sequence and Tensor Parallelism
Memory Orchestration
Hierarchical Parallel Training
Activation Offloading
High-Resolution Earth Forecasting
🔎 Similar Papers
No similar papers found.
R
Ruohan Wu
School of Computer Science and Technology, University of Science and Technology of China, Hefei 230026, China
Ziqi Zhu
Ziqi Zhu
PhD student, University of Science and Technology of China
AI for ScienceComputer Vision
Y
Yang Zhao
School of Computer Science and Technology, University of Science and Technology of China, Hefei 230026, China; and Department of Ocean Big Data and Artificial Intelligence, Laoshan Laboratory, Qingdao, China
J
Jiarui Tang
School of Computer Science and Technology, University of Science and Technology of China, Hefei 230026, China; and Department of Ocean Big Data and Artificial Intelligence, Laoshan Laboratory, Qingdao, China
Y
Yingzhe Cui
Department of Ocean Big Data and Artificial Intelligence, Laoshan Laboratory, Qingdao, China
J
Junshi Chen
School of Computer Science and Technology, University of Science and Technology of China, Hefei 230026, China; and Department of Ocean Big Data and Artificial Intelligence, Laoshan Laboratory, Qingdao, China
Z
Zhao Jing
Frontiers Science Center for Deep Ocean Multispheres and Earth System and Key Laboratory of Physical Oceanography/Academy of Future Ocean, Ocean University of China, Qingdao, China; and Department of Ocean Big Data and Artificial Intelligence, Laoshan Laboratory, Qingdao, China
Jun Shi
Jun Shi
Post Doc, University of Science and Technology of China
Medical Image AnalysisAI for ScienceInference & Training Optimization
H
Hong An
School of Computer Science and Technology, University of Science and Technology of China, Hefei 230026, China; and Department of Ocean Big Data and Artificial Intelligence, Laoshan Laboratory, Qingdao, China