Characterizing the Scalability and Performance of Large-Scale AI Training Under Multi-Tenancy

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过评估不同并行策略和配置下的性能,解决了大规模AI训练在多租户环境下的可扩展性和执行效率问题。
📝 Abstract
Characterising AI workload performance on modern HPC systems requires understanding both their scalability in isolation and their behaviour under concurrent execution. However, the interplay among parallelisation strategies, network congestion, compute capability, and interconnect technologies remains poorly understood. This work investigates the performance and scalability of AI models up to 2400 GPUs. We quantify the communication overheads and their impact across different interconnects by evaluating scale-up, scale-out, and rack-scale configurations under multiple allocation schemes. Finally, we study how multiple concurrent training jobs interfere with each other by designing a realistic noise model. We design a benchmark suite of AI models to evaluate the performance of five distinct parallelisation strategies across different supercomputing clusters, including Alps, Leonardo, LUMI, JUPITER, NVL72 GB300, and DGX A100. Our work provides a systematic characterization of the scalability and execution efficiency of distributed AI training, while offering key insights into performance behavior under realistic multi-tenant scenarios.
Problem

Research questions and friction points this paper is trying to address.

Scalability
Performance
Multi-Tenancy
AI Training
Interconnect Technologies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scalability
Multi-tenancy
Communication Overheads
Concurrent Training
Interconnect Technologies
🔎 Similar Papers
No similar papers found.