MatrixFlow: System-Accelerator co-design for high-performance transformer applications
Transformer models face severe acceleration bottlenecks due to their high computational demands and memory bandwidth requirements. Method: This paper proposes MatrixFlow, a hardware–software co-designed architecture featuring a novel loosely coupled systolic array and a dataflow-driven matrix multiplication mechanism, enabling system-level joint optimization of computation, data movement, and memory access. It further introduces a flexible hardware–software mapping algorithm supporting diverse models—including BERT and ViT—and validates the design via full-system gem5 simulation. Contribution/Results: Experiments show that MatrixFlow achieves up to 22× speedup over many-core CPUs, and outperforms state-of-the-art loosely coupled and tightly coupled accelerators by 5× and 8×, respectively. It significantly reduces memory overhead and improves energy efficiency.