Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决AI计算中的效率与扩展性问题,本文介绍了一种名为Maia 200的软件定义数据流系统,通过优化数据流动和内存管理来提高性能并减少能耗。
📝 Abstract
We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth. Maia exemplifies a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), which explicitly program dataflow engines to orchestrate highly specialized memories and data movement engines. This approach shifts the focus from today's thread-centric to data-movement-centric architecture, improving efficiency and scalability. Our taxonomy of data management, inspired by Flynn's classification, highlights how SDLA addresses challenges in modern AI computing. Maia 200 achieves significant cost and energy savings while supporting massive parallelism for AI inference workloads, making it a compelling solution for next-generation high-performance computing systems.
Problem

Research questions and friction points this paper is trying to address.

AI Acceleration
Data Movement
Scalability
Energy Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Software Defined Locally Accessed Dataflow Architectures
data-movement-centric architecture
massive parallelism
🔎 Similar Papers
No similar papers found.
S
Sherry Xu
Microsoft Corporation
M
Marco Heddes
Microsoft Corporation
J
Jackson Peng
Microsoft Corporation
T
Tom Savell
Microsoft Corporation
M
Monica Tang
Microsoft Corporation
P
Prashant Ranjan
Microsoft Corporation
J
Jesse Benson
Microsoft Corporation
O
Ofer Dekel
Microsoft Corporation
S
Saurabh Dighe
Microsoft Corporation
A
Anupama Kurpad
Microsoft Corporation
A
Artour Levin
Microsoft Corporation
Matthew Mattina
Matthew Mattina
Microsoft Corporation
G
George Petre
Microsoft Corporation
C
Cheng Tang
Microsoft Corporation
Y
Yuan Yu
Microsoft Corporation
Li Zhang
Li Zhang
Microsoft
algorithms
Torsten Hoefler
Torsten Hoefler
Professor of Computer Science at ETH Zurich
High Performance ComputingDeep LearningNetworkingMessage Passing InterfaceParallel and Distributed Computing