CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference Acceleration

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
CHIPSMORE通过集成计算互连和内存计算技术,解决大语言模型推理在不同模式下的高利用率、内存效率和性能扩展问题。
📝 Abstract
Large language model (LLM) inference exhibits substantial variability across adaptation modes, context lengths, and request concurrency, creating challenges for maintaining high utilization, memory efficiency, and scalable performance on compute-in-memory (CIM) accelerators. This paper presents CHIPSMORE, a multi-mode and multi-request LLM inference accelerator that integrates compute-in-interconnect and CIM to support both base-mode and low-rank adaptation (LoRA) inference under diverse workloads. CHIPSMORE employs heterogeneous processing elements consisting of resistive RAM analog compute-in-memory (RRAM-ACIM) and static RAM digital compute-in-memory (SRAM-DCIM) interconnected through a programmable Inter-PE computational network (IPCN). A composable hierarchical key-value (KV) memory scheme dynamically allocates router scratchpad, SRAM-DCIM, and embedded DRAM (eDRAM) resources according to workload requirements, enabling scalable support for long-context and batched inference. Furthermore, a non-replicated multi-request execution pipeline exploits request-level parallelism without duplicating pretrained weights, while a state-aware resource reconfiguration mechanism selectively retains runtime states and power-gates inactive resources to improve energy efficiency. Evaluation using cycle-accurate hardware-software co-simulation demonstrates that CHIPSMORE effectively sustains high throughput across varying model sizes, context lengths, and batch sizes while maintaining favorable power scaling. Compared with Nvidia H100, CHIPSMORE achieves up to $2.38\times$ higher throughput and $27\times$ higher energy efficiency on Mistral-7B inference while eliminating weight replication for multi-request serving.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model
Inference Variability
Compute-in-Memory Accelerators
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compute-in-Interconnect
Multi-Mode Inference
Non-replicated Multi-Request Pipeline
Hierarchical Key-Value Memory
State-Aware Resource Reconfiguration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yue Jiet Chong
Department of Electrical and Computer Engineering, National University of Singapore, Singapore
Yimin Wang
Yimin Wang
National University of Singapore | Fudan University | Jilin University
Circuits and SystemsIn-Memory ComputingAI AcceleratorHardware/Software Co-Design
Z
Zhen Wu
Department of Electrical and Computer Engineering, National University of Singapore, Singapore
Z
Zixuan Wang
Department of Electrical and Computer Engineering, National University of Singapore, Singapore
W
Wei Zhang
Department of Electrical and Computer Engineering, National University of Singapore, Singapore
X
Xuanyao Fong
Department of Electrical and Computer Engineering, National University of Singapore, Singapore