An HPC Approach to Accelerate Tensor Decompositions

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对高维数据处理难题,提出一种新的基于CUDA的Jacobi型张量分解算法,并在NVIDIA H100 GPU上验证其有效性及高效性。
📝 Abstract
Quantum systems grow in complexity so rapidly that even modest models become difficult to simulate, creating a strong need for methods that can handle high-dimensional data, also known as tensors. In this work, we investigate a novel Jacobi-type tensor algorithm for tensor decomposition and develop a CUDA-based algorithm that supports tensors of arbitrary order on a single GPU. We test the implementation on NVIDIA H100 GPUs and show that the algorithm converges correctly for diagonalizable tensors up to nine dimensions, with runtime scaling in a predictable way as tensor order grows. Finally, our general algorithm outperforms the original MATLAB reference by more than two orders of magnitude.
Problem

Research questions and friction points this paper is trying to address.

high-dimensional data
tensors
quantum systems
simulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

CUDA-based algorithm
arbitrary order tensors
predictable runtime scaling
two orders of magnitude performance improvement
M
Markus Hellgren
Uppsala University, Department of Information Technology, Sweden
E
Erna Begovic Kovac
University of Zagreb, Faculty of Chemical Engineering and Technology, Croatia
H
Hans O. Karlsson
Uppsala University, Department of Information Technology, Sweden
Roman Iakymchuk
Roman Iakymchuk
Uppsala University and Umeå University
Energy-efficient computingreliable computinghigh-performance computingnumerical linear algebra