Uni-SFU: Algorithm-HW Co-Design for Universal SFUs via Mixed-Degree Piecewise Approximation

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of balancing computational efficiency, area overhead, and numerical accuracy in hardware implementations of activation functions. The authors propose a novel algorithm-hardware co-design framework that introduces, for the first time, a cross-activation mixed-order, non-uniform piecewise polynomial approximation strategy. This approach is jointly optimized with an RTL-level area cost model to achieve high accuracy and hardware reuse under a unified configuration. Experimental results demonstrate that the proposed design attains a mean squared error below 8.22×10⁻⁸ and incurs no more than a 1.02% Top-1 accuracy loss across over 700 neural network variants and three NLP models. Implemented in 22 nm technology, the hardware footprint is only 6,800 μm².
📝 Abstract
Nonlinear activation functions are essential to modern deep neural networks (DNNs), but their hardware evaluation places significant pressure on the special-function units (SFUs) of GPUs and custom accelerators. Therefore, piecewise polynomial approximations are commonly used within allowed error bounds to improve computational efficiency. However, existing techniques often approximate each activation function in isolation using fixed-degree polynomials and uniform segments, leading to hardware redundancy and sub-optimal precision. To address these limitations, we present Uni-SFU, an algorithm-hardware co-design framework that jointly optimizes approximation accuracy and silicon area for a diverse set of activation functions. Uni-SFU leverages a joint search across all target functions to assign mixed-degree polynomials to nonuniform segments, guided by an RTL-derived area cost model. This approach identifies a unified hardware configuration to implement the target activation functions under given accuracy constraints. Validated across over 700 neural network variants and three Natural Language Processing (NLP) models, Uni-SFU achieves a superior Mean Squared Error (MSE) below 8.22x10^-8, limiting top-1 accuracy degradation to within 1.02% compared to floating-point baselines. The proposed design occupies only 6,800 um2 in GF 22nm CMOS technology, achieving a superior trade-off between silicon area and system-level accuracy compared to SOTA counterparts.
Problem

Research questions and friction points this paper is trying to address.

activation functions
special-function units
piecewise polynomial approximation
hardware redundancy
approximation accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

algorithm-hardware co-design
mixed-degree piecewise approximation
special-function units (SFUs)
nonlinear activation functions
area-accuracy trade-off
🔎 Similar Papers
No similar papers found.
Miao Sun
Miao Sun
WeRide
Computer VisionAutonomous Driving
Y
Yucheng Huang
Department of Electrical and Computer Engineering, University of Wisconsin–Madison, Madison, WI 53706 USA
M
Mingcong Cao
Department of Electrical and Computer Engineering, University of Wisconsin–Madison, Madison, WI 53706 USA
Jaehyun Park
Jaehyun Park
Assistant Professor, School of Electrical Engineering, University of Ulsan
Low-power designIoT system design
P
Partha Pratim Pande
School of Electrical Engineering and Computer Science, Washington State University, Pullman, WA 99164 USA
U
Umit Y. Ogras
Department of Electrical and Computer Engineering, University of Wisconsin–Madison, Madison, WI 53706 USA