Institution profile

Laoshan Laboratory

Academic institutionasia · cn
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

TERRA: A Hierarchical Parallel Training and Memory Orchestration Framework for High-Resolution AI-based Earth Modeling

Aug 15, 2026

This study addresses GPU memory bottlenecks and inadequate distributed system support in training high-resolution AI Earth prediction models by proposing a hierarchical parallel training framework. The approach employs a sampling window sequence-aware tensor parallelism strategy to preserve spatial continuity, integrated with rollout-aware checkpointing and budget-constrained activation offloading to optimize memory orchestration. This method effectively resolves computational and storage challenges in long-sequence fine-tuning. Evaluated on a 96-GPU H200 cluster, the framework supports training an 11.4-billion-parameter model, achieving a peak performance of 39.76 PFLOPS with strong and weak scaling efficiencies of 65% and 94.1%, respectively, while reducing peak GPU memory usage by over 32%.

0 citationsRead paper
Recent publications

Latest Papers

TERRA: A Hierarchical Parallel Training and Memory Orchestration Framework for High-Resolution AI-based Earth Modeling

Aug 15, 2026

This study addresses GPU memory bottlenecks and inadequate distributed system support in training high-resolution AI Earth prediction models by proposing a hierarchical parallel training framework. The approach employs a sampling window sequence-aware tensor parallelism strategy to preserve spatial continuity, integrated with rollout-aware checkpointing and budget-constrained activation offloading to optimize memory orchestration. This method effectively resolves computational and storage challenges in long-sequence fine-tuning. Evaluated on a 96-GPU H200 cluster, the framework supports training an 11.4-billion-parameter model, achieving a peak performance of 39.76 PFLOPS with strong and weak scaling efficiencies of 65% and 94.1%, respectively, while reducing peak GPU memory usage by over 32%.

0 citationsRead paper