LLM Inference on IMC-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种硬件-软件协同设计框架,通过整合IMC PE、NMC PE和INC来解决大规模语言模型推理中的内存带宽和计算密度问题,提高吞吐量和能效。
📝 Abstract
LLM inference has become an essential service, yet it imposes unprecedented demands on memory bandwidth, computational density, and communication efficiency. While IMC is a promising solution to the memory wall issue, the heterogeneous data dynamicity of LLM requires complementary resources to handle intermediate data generated during run-time. Furthermore, the massive number of parameters in LLM necessitates scale-up architectures where on-chip data movement is often the primary performance bottleneck. This article presents a hardware-software co-design framework that unifies distributed compute, memory, and communication into a seamless processing-communication fabric. On the hardware side, we propose a scalable architecture, named LEAP, that integrates IMC PE, NMC PE, and INC. This allows each hardware layer to execute specialized tasks: IMC for static weights, NMC for dynamic data, and INC for partial result reduction. On the software side, we introduce a partitioning, mapping, and scheduling framework optimized for key metrics in LLM serving, including throughput and latency. To address the distinct computational intensities of the prefill and decode phases, we present a prefill-decode disaggregation approach that dynamically reconfigures PE organizations to maximize resource utilization. Compared to commercial GPU platforms, the proposed architecture provides a throughput and an energy efficiency improvement of $\geq{}1.52\times$ and $24.91\times$, respectively.
Problem

Research questions and friction points this paper is trying to address.

LLM Inference
Memory Bandwidth
Computational Density
Communication Efficiency
On-chip Data Movement
Innovation

Methods, ideas, or system contributions that make the work stand out.

hardware-software co-design
LEAP architecture
prefill-decode disaggregation
throughput improvement
energy efficiency
🔎 Similar Papers
No similar papers found.