RACE-IT: A Reconfigurable Analog CAM-Crossbar Engine for In-Memory Transformer Acceleration

📅 2023-11-29
🏛️ arXiv.org
📈 Citations: 5
Influential: 2
📄 PDF
🤖 AI Summary
Transformer inference on in-memory computing (IMC) platforms faces bottlenecks due to inefficient execution of non-vector-matrix-multiplication (non-VMM) operations—particularly Softmax and data-dependent matrix multiplications. This work proposes Compute-ACAM, a reconfigurable analog-domain compute architecture tightly integrated with crossbar arrays, featuring the first analog-domain content-addressable compute unit capable of executing arbitrary non-VMM operations—enabling end-to-end analog-domain Transformer inference. It introduces an ADC-free quantized output circuit that directly converts analog inputs to digital outputs. The hardware is natively compatible with emerging DNN architectures without structural modification. Compared to state-of-the-art GPUs and IMC accelerators, Compute-ACAM achieves 10.7× and 5.9× higher throughput, and reduces energy consumption by 1193× and 3.9×, respectively—significantly overcoming IMC’s scalability limitations for complex DNN operators.
📝 Abstract
Transformer models represent the cutting edge of Deep Neural Networks (DNNs) and excel in a wide range of machine learning tasks. However, processing these models demands significant computational resources and results in a substantial memory footprint. While In-memory Computing (IMC) offers promise for accelerating Matrix-Vector Multiplications (MVMs) with high computational parallelism and minimal data movement, employing it for implementing other crucial operators within DNNs remains a formidable task. This challenge is exacerbated by the extensive use of Softmax and data-dependent matrix multiplications within the attention mechanism. Furthermore, existing IMC designs encounter difficulties in fully harnessing the benefits of analog MVM acceleration due to the area and energy-intensive nature of Analog-to-Digital Converters (ADCs). To tackle these challenges, we introduce a novel Compute Analog Content Addressable Memory (Compute-ACAM) structure capable of performing various non-MVM operations within Transformers. Together with the crossbar structure, our proposed RACE-IT accelerator enables efficient execution of all operations within Transformer models in the analog domain. Given the flexibility of our proposed Compute-ACAMs to perform arbitrary operations, RACE-IT exhibits adaptability to diverse non-traditional and future DNN architectures without necessitating hardware modifications. Leveraging the capability of Compute-ACAMs to process analog input and produce digital output, we also replace ADCs, thereby reducing the overall area and energy costs. By evaluating various Transformer models against state-of-the-art GPUs and existing IMC accelerators, RACE-IT increases performance by 10.7x and 5.9x, and reduces energy by 1193x, and 3.9x, respectively
Problem

Research questions and friction points this paper is trying to address.

Accelerating Transformer models with in-memory computing
Supporting complex DNN operations in analog domain
Enhancing flexibility for emerging DNN architectures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reconfigurable Analog CAM-Crossbar Engine for acceleration
Enhances Analog Content Addressable Memories for broader operations
Supports analog-domain execution of Transformer core operations
Hewlett Packard Labs
L
Lei Zhao
Artificial Intelligence Research Lab (AIRL), Hewlett Packard Labs, USA
L
Luca Buonanno
Artificial Intelligence Research Lab (AIRL), Hewlett Packard Labs, USA
R
Ron M. Roth
Artificial Intelligence Research Lab (AIRL), Hewlett Packard Labs, USA
S
S. Serebryakov
Artificial Intelligence Research Lab (AIRL), Hewlett Packard Labs, USA
Archit Gajjar
Archit Gajjar
Hewlett Packard Labs, HPE
Hardware AccelerationFPGAsMachine Learning
J
J. Moon
Artificial Intelligence Research Lab (AIRL), Hewlett Packard Labs, USA
J
J. Ignowski
Artificial Intelligence Research Lab (AIRL), Hewlett Packard Labs, USA
Giacomo Pedretti
Giacomo Pedretti
Research Scientist, Hewlett Packard Laboratories
AI acceleratorsIn-memory computingNeuromorphic ComputingAnalog computingEmerging memories