DB-SpMSpV: Dual-View Blocked Sparse Matrix-Sparse Vector Multiplication for Dynamic GPU Workloads

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses performance bottlenecks in GPU dynamic sparse matrix-vector multiplication caused by varying input sparsity. We propose a dual-view tiling framework that employs 2D tiled storage with shared underlying data in both CSR and CSC formats to support push-pull traversal. The method further optimizes dynamic workloads through runtime adaptive kernel selection, load balancing, and asynchronous prefetching. Experiments on NVIDIA A100 GPUs demonstrate substantial improvements over cuSPARSE, achieving average speedups ranging from 5.48× to 64.34×. Additionally, the framework accelerates graph traversal by 2.66× and linear layers in large model decoding by 4.50×, significantly enhancing computational efficiency for dynamic sparse workloads.
📝 Abstract
Sparse Matrix-Sparse Vector Multiplication (SpMSpV) is a core primitive in graph traversal, sparse linear algebra, and sparse model inference. Its input vector is often dynamically sparse, so the best GPU execution path depends on both global sparsity and the local vector-block distribution. Existing GPU SpMSpV methods often bind storage layouts, push/pull traversal, and kernels together, making fine-grained adaptation difficult without extra storage or scheduling overhead. This paper presents DB-SpMSpV, a dual-view blocked SpMSpV framework for dynamic GPU workloads. DB-SpMSpV partitions the matrix into fixed-size 2D blocks, maintains block-level CSR/CSC views at the high level, and reuses a single low-level block payload to support both row-driven pull and column-driven push. At runtime, it selects the global traversal path based on input block sparsity, chooses block microkernels from the local matrix/vector block structure, and uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance. We further integrate the framework into DB-BFS and DB-Decoding. We evaluate DB-SpMSpV on NVIDIA A100 and RTX 4090 using SuiteSparse matrices, symmetric graphs, and three open-source LLMs. Across input sparsities, DB-SpMSpV achieves average speedups of 5.48$\times$--64.34$\times$ over cuSPARSE and 2.36$\times$--14.01$\times$ over TileSpMSpV on A100, with similar gains on RTX 4090. DB-BFS further improves end-to-end graph traversal by 2.66$\times$ over TileBFS on A100 and 3.60$\times$ on RTX 4090 on average, while DB-Decoding accelerates single-token linear layers by up to 4.50$\times$.
Problem

Research questions and friction points this paper is trying to address.

Sparse Matrix-Sparse Vector Multiplication
Dynamic GPU Workloads
Fine-grained Adaptation
Storage Layout Coupling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-View Blocked SpMSpV
Dynamic GPU Workloads
Adaptive Traversal
Storage-Traversing Decoupling
Hierarchical Writeback
🔎 Similar Papers
No similar papers found.