Hardware Acceleration of Block-Diffusion LLM for Edge Devices

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了边缘设备上单流推理权重流量无法分摊的问题,通过设计WIFiV-LPDDR、BRQ-KV和DAT-FFN方法优化了块扩散LLM模型的效率。
📝 Abstract
Single-stream (batch-one) edge inference cannot amortize weight traffic across requests. Full-attention diffusion LLMs recompute the entire sequence at every step; native block diffusion makes completed blocks immutable and exactly cacheable, yet refinement still streams prefix KV and FFN weights. We co-design WIFiV-LPDDR, a wide-I/O LPDDR system for precision-tagged reads, BRQ-KV for a canonical low-rank-plus-INT8-residual prefix with query-dependent per-entry precision, and DAT-FFN for drift-mapped canonical replacement, adjacent-stage-corrected low-bit delta, or cached-state carry while keeping live activations unquantized. Both map to an input-stationary mixed-precision systolic array. For the evaluated 1.5B/7B models on modeled Jetson-class platforms, the full stack provides arithmetic-mean energy-reduction factors of 3.79x/3.96x and arithmetic-mean latency speedups of 2.88x/4.44x at the reported DAT-FFN settings; every corresponding compressed model-benchmark score drops by less than one absolute percentage point from its baseline.
Problem

Research questions and friction points this paper is trying to address.

Hardware Acceleration
Block-Diffusion LLM
Edge Devices
Innovation

Methods, ideas, or system contributions that make the work stand out.

WIFiV-LPDDR
BRQ-KV
DAT-FFN
mixed-precision systolic array
W
Wei-Hsing Huang
School of Electrical and Computer Engineering, Georgia Institute of Technology, Atlanta, GA 30332 USA
K
Kiseok Lee
School of Electrical and Computer Engineering, Georgia Institute of Technology, Atlanta, GA 30332 USA
Ming-Yen Lee
Ming-Yen Lee
ECE, Georgia Institute of Technology
In-memory ComputingIntegrated Circuits and SystemsSoftware-Hardware Codesign
W
Weiyu Sun
School of Computer Science, Georgia Institute of Technology, Atlanta, GA 30332 USA
Cheng-Jhih Shih
Cheng-Jhih Shih
National Taiwan University
High-performance computingHomomorphic encryptionQuantum computingPost quantum cryptography
G
Gayatri Tanksali
School of Electrical and Computer Engineering, Georgia Institute of Technology, Atlanta, GA 30332 USA
A
Arpit Khandelwal
School of Computer Science, Georgia Institute of Technology, Atlanta, GA 30332 USA
P
Pin-Jun Chen
School of Electrical and Computer Engineering, Georgia Institute of Technology, Atlanta, GA 30332 USA
Yingyan (Celine) Lin
Yingyan (Celine) Lin
Associate Professor, Georgia Institute of Technology
Efficient AI algorithmsDeep learning acceleratorsGreen AI
Shimeng Yu
Shimeng Yu
Georgia Institute of Technology, Dean's Professor
Non-volatile MemoryRRAMFerroelectric MemoriesIn-Memory ComputingAI Hardware