REACH: Controller-Managed Long-Span ECC for HBM AI Inference

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决HBM在AI推理中的高成本问题,REACH通过结合短码和长码的错误校正策略,在保持高带宽的同时降低了控制器面积和功耗。
📝 Abstract
High-Bandwidth Memory (HBM) cost motivates stronger controller protection that can support a wider range of device error rates. Long-span error-correcting codes provide stronger protection at a comparable code rate, but a direct implementation couples small accesses to span-wide state and requires costly decoding at HBM bandwidth. Read-dominated LLM decode offers a favorable setting: sequential reads support span aggregation, while sparse writes limit parity-update traffic. This paper presents REACH, a controller microarchitecture that uses established inner codes to correct common errors and identify unresolved chunks, reserving a long outer code for known-erasure repair. Differential parity bounds write traffic, and a co-designed endpoint preserves 32\,B transactions without an extra data burst. Ramulator2 sustains 1.88\,TB/s of application traffic at the highest error stress, while separate full-interface sizing supports a 2.69\,TB/s application target using ASAP7-synthesized kernels. At this analytical target, REACH's nominal composition uses 55.8\% less controller area and 57.7\% less modeled power than the evaluated mean-work direct-long design, showing the benefit of reserving long-span recovery for exceptional requests.
Problem

Research questions and friction points this paper is trying to address.

High-Bandwidth Memory
Controller Protection
Error-Correcting Codes
Error Rates
Innovation

Methods, ideas, or system contributions that make the work stand out.

controller-managed
long-span ECC
HBM AI inference
differential parity
co-designed endpoint
🔎 Similar Papers
No similar papers found.