Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Para-Pipe框架,通过结合操作符并行和流水线技术优化SoC上深度学习应用的吞吐量与延迟,减少处理器间通信开销,提高能效。
📝 Abstract
As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. Traditional pipelining techniques distributing the computation across different on-chip processing units, while effective for throughput, do not address the latency demands posed by modern neural networks with complex interdependencies and extensive operator parallelism. There is a potential in leveraging operator parallelism to enable concurrent execution across multiple processing units, thereby reducing inference latency. However, prioritizing pipelining or parallel execution often necessitates a compromise, where optimizing one performance metric adversely impacts the other. This paper introduces Para-Pipe, a hierarchical mapping framework that integrates intra- and inter-stage operator parallelism within a pipelined architecture. Para-Pipe navigates the trade-off between throughput and latency by selectively fine-tuning parallelism levels within and across pipeline stages. This strategy can significantly reduce inter-processor communication overhead, significantly improving energy efficiency. Our evaluation demonstrates that Para-Pipe generates multiple Pareto-optimal configurations, achieving a balance between throughput and latency on an Amlogic SoC equipped with ARM big.LITTLE CPUs and GPU, as well as the Black Sesame Technology SoC featuring a deep learning accelerator and two DSPs. More importantly, throughput-optimized configurations under Para-Pipe on Amlogic SoC show an average energy efficiency improvement of 11.0% over purely pipelined strategies and 23.3% relative to non-pipelined parallel execution.
Problem

Research questions and friction points this paper is trying to address.

heterogeneous System-on-Chips
operator parallelism
inference latency
throughput
energy efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Operator Parallelism
Pipelined Architecture
Throughput and Latency Trade-off
Energy Efficiency
🔎 Similar Papers
No similar papers found.
Yujie Zhang
Yujie Zhang
Shanghai Jiao tong University
3D Quality AssessmentGeometry Processing3D Reconstruction
H
Huiying Lan
School of Computing, National University of Singapore, Singapore
E
Ehsan Aghapour
Parallel Computer Systems, University of Amsterdam, 1098 XH Amsterdam, The Netherlands
Zhiyuan Ning
Zhiyuan Ning
Westlake University
Graph Machine LearningKnowledge GraphsLarge Language Models
P
Peng Zan
Information Technology Department, Black Sesame Technologies, San Jose, CA 95131, USA
W
Weidong Shao
Information Technology Department, Black Sesame Technologies, San Jose, CA 95131, USA
Anuj Pathania
Anuj Pathania
Assistant Professor, University of Amsterdam
Electronic Design Automation
Tulika Mitra
Tulika Mitra
Professor of Computer Science, National University of Singapore
Design AutomationLow Power DesignEmbedded SystemsReal-Time Systems