Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs
本文提出Para-Pipe框架,通过结合操作符并行和流水线技术优化SoC上深度学习应用的吞吐量与延迟,减少处理器间通信开销,提高能效。
本文提出Para-Pipe框架,通过结合操作符并行和流水线技术优化SoC上深度学习应用的吞吐量与延迟,减少处理器间通信开销,提高能效。
This work addresses the challenge of reliably tracing the provenance of tensors and operators through graph rewrites—particularly non-injective transformations—in AI compilers. The authors propose a lightweight, generative provenance method grounded in observational semantics, which infers origins by analyzing the behavioral effects of graph transformations rather than relying on identifier propagation. For the first time, they introduce coalgebraic modeling and bisimulation to this domain, guaranteeing provenance consistency even after intermediate nodes are eliminated. The approach requires no invasive compiler modifications and naturally supports non-injective rewrites. Evaluated within COVAN, a prototype AI compiler, the method demonstrates stable, low-overhead provenance tracking throughout an end-to-end compilation pipeline.
This work addresses the insufficient visual constraints and representational redundancy in perception-agnostic end-to-end autonomous driving, where scene tokens are supervised solely by planning objectives. To mitigate this, the authors propose a Neural Token Reconstruction (NTR) framework that introduces, for the first time, a self-distilled masked latent reconstruction objective at the scene token bottleneck. This approach leverages compact tokens as memory to reconstruct patch-level image features, thereby enhancing their representational capacity. Guided by semantic priors from foundation models, the reconstruction process focuses on driving-relevant structures without requiring additional modules during inference. Evaluated on Waymo E2E and NavSim1&2 benchmarks, NTR achieves state-of-the-art performance (RFS: 8.0461; PDMS/EPDMS: 94.1/90.9), significantly reducing token redundancy and improving effective rank.
本文提出Para-Pipe框架,通过结合操作符并行和流水线技术优化SoC上深度学习应用的吞吐量与延迟,减少处理器间通信开销,提高能效。
This work addresses the challenge of reliably tracing the provenance of tensors and operators through graph rewrites—particularly non-injective transformations—in AI compilers. The authors propose a lightweight, generative provenance method grounded in observational semantics, which infers origins by analyzing the behavioral effects of graph transformations rather than relying on identifier propagation. For the first time, they introduce coalgebraic modeling and bisimulation to this domain, guaranteeing provenance consistency even after intermediate nodes are eliminated. The approach requires no invasive compiler modifications and naturally supports non-injective rewrites. Evaluated within COVAN, a prototype AI compiler, the method demonstrates stable, low-overhead provenance tracking throughout an end-to-end compilation pipeline.
This work addresses the insufficient visual constraints and representational redundancy in perception-agnostic end-to-end autonomous driving, where scene tokens are supervised solely by planning objectives. To mitigate this, the authors propose a Neural Token Reconstruction (NTR) framework that introduces, for the first time, a self-distilled masked latent reconstruction objective at the scene token bottleneck. This approach leverages compact tokens as memory to reconstruct patch-level image features, thereby enhancing their representational capacity. Guided by semantic priors from foundation models, the reconstruction process focuses on driving-relevant structures without requiring additional modules during inference. Evaluated on Waymo E2E and NavSim1&2 benchmarks, NTR achieves state-of-the-art performance (RFS: 8.0461; PDMS/EPDMS: 94.1/90.9), significantly reducing token redundancy and improving effective rank.