The Tensor-Core Beamformer: A High-Speed Signal-Processing Library for Multidisciplinary Use
Beamforming in multi-sensor signal processing suffers from high computational density, poor hardware adaptability, and insufficient flexibility in precision. Method: This paper introduces the first tensor-core-oriented general-purpose beamforming acceleration library. It pioneers the generalization of GPU tensor cores for beamforming computation, integrates mixed-precision (FP16/1-bit) design, and achieves cross-platform high-efficiency deployment on both NVIDIA and AMD GPUs via dual-stack heterogeneous optimization using CUDA and HIP. Contributions/Results: The library achieves over 600 TeraOps/s (FP16) on AMD MI300X with near 1 TeraOp/J energy efficiency; on NVIDIA A100, it delivers 3 PetaOps/s in 1-bit mode with >10 TeraOps/J efficiency. It enables, for the first time, ultra-low-precision real-time beamforming and has been successfully deployed in clinical ultrasound and radio astronomy systems.