Fengshui: Demystifying Chiplet Ecosystem and Bespoke Neural Network Accelerator Codesign

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Fengshui框架,通过芯片粒生态系统与定制ASIC协同设计,解决现代机器学习工作负载在通用硬件上运行效率低的问题。
📝 Abstract
Modern ML workloads, with stringent latency and energy constraints, are increasingly hard to run efficiently on homogeneous commodity hardware. We argue that operator-level disaggregation--tailoring microarchitecture, batching, and memory hierarchy to each operator--is essential to overcome these limitations, though the resulting highly bespoke accelerators incur prohibitive Non-Recurring Engineering (NRE) costs. Chiplet-based integration amortizes NRE across applications, but choosing which chiplets to build and how to compose them into accelerators is circularly dependent--a chiplet pool's value depends on the constructed accelerators, while accelerator quality is constrained by available chiplets. This paper introduces Fengshui, a chiplet ecosystem and accelerator co-design framework that jointly optimizes chiplet pool composition and bespoke application-specific integrated circuit (BASIC) design. Fengshui constructs BASICs through operator-level disaggregation, co-exploring chiplet and memory heterogeneity, tensor fusion, and pipeline/tensor/expert parallelism with place-and-route validation for physical implementability. With just 8 strategically selected chiplets, encompassing network switches, processing-in-memory units, and accelerators with diverse microarchitectures, Fengshui-generated BASICs achieve 48.5%, 88.1%, 93.0%, and 97.8% reductions in energy, energy-cost product (EC), energy-delay product (EDP), and energy-delay-cost product (EDPC) over homogeneous accelerators, while scoring within 4.1% of unconstrained heterogeneous designs across diverse neural networks. For datacenter MoE and dense LLM serving, Fengshui reduces prefill energy and EC by up to 16.8% and 28.7%, respectively; for edge autonomous vehicle perception, it achieves 12.0% energy and 23.6% EC reductions under real-time latency constraints.
Problem

Research questions and friction points this paper is trying to address.

Machine Learning Workloads
Latency and Energy Constraints
Operator-level Disaggregation
Non-Recurring Engineering (NRE) Costs
Chiplet-based Integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chiplet Ecosystem
Operator-Level Disaggregation
Bespoke Accelerator Design
Heterogeneous Integration
Energy Efficiency
🔎 Similar Papers
No similar papers found.
Haoran Jin
Haoran Jin
Zhejiang University
acousticultrasoundphotoacoustic
J
Jirong Yang
Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor, MI, USA
Z
Zhiheng Zhang
Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor, MI, USA
J
Justin Shin
Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor, MI, USA
B
Barry Lyu
Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor, MI, USA
K
Kangqi Zhang
Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor, MI, USA
Yunpeng Liu
Yunpeng Liu
Wuhan University of Technology
cement and concrete materials
N
Nathan Bleier
Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor, MI, USA