ScaleLUT: A Fully-Parallel Configurable LUT-Based Accelerator for Real-Time Multi-Scale Super-Resolution

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对边缘设备实时超分辨率挑战,提出ScaleLUT,一种基于查找表的全并行可配置加速器,通过硬件友好设计减少计算和存储需求。
📝 Abstract
Real-time super-resolution (SR) remains challenging for edge devices because deep-learning-based methods require substantial multiply-accumulate (MAC) operations, resources, and power. Lookup-table (LUT)-based SR reduces computation by replacing convolutional inference with table queries, but existing methods still suffer from limited speed, large storage overhead, and poor scalability across upsampling factors. We present ScaleLUT, a hardware-oriented LUT design framework and fully parallel reconfigurable accelerator for real-time multi-scale SR. ScaleLUT combines a hardware-friendly YUV-domain strategy with power-of-two kernels and rotation ensemble to improve receptive-field coverage while reducing LUT dimensionality; division operations are replaced by shifts. These designs reduce memory by 18.4% over state-of-the-art LUT-based SR methods. ScaleLUT supports arbitrary input resolutions and configurable x2^n upsampling factors using a deeply pipelined, massively parallel architecture. Implemented on a Xilinx ZCU102 FPGA, it achieves real-time 4K SR at 95.3 FPS for x2 upscaling at 300 MHz. Compared with existing SR accelerators, ScaleLUT uses at least 58.6% fewer LUTs, 41.1% fewer flip-flops, zero DSPs, and 42.0% lower power, while delivering 10x and 1.2x speedups over the best CPU-based SR implementation and prior FPGA-based SR accelerators, respectively. These results demonstrate the effectiveness of joint LUT algorithm-hardware co-design for practical and energy-efficient edge SR deployment.
Problem

Research questions and friction points this paper is trying to address.

Real-time super-resolution
Edge devices
Multiply-accumulate operations
Lookup-table
Scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fully-Parallel Configurable LUT
YUV-domain strategy
Power-of-two kernels
Rotation ensemble
Real-time multi-scale SR
🔎 Similar Papers
No similar papers found.
Boyu Li
Boyu Li
Hong Kong University of Science and Technology
3D interactionHuman-AI Collaboration
C
Chenchen Ding
The University of Hong Kong (HKU), Hong Kong
Z
Zhilin Ai
The University of Hong Kong (HKU), Hong Kong
W
Wenqing Shi
The University of Hong Kong (HKU), Hong Kong
B
Baizhou Jiang
The University of Hong Kong (HKU), Hong Kong
Wenyong Zhou
Wenyong Zhou
The University of Hong Kong
Computer Vision
Binxiao Huang
Binxiao Huang
PhD Candidate, HKU
trustworthy and efficient AI3D vision
J
Jiachen Ren
The University of Hong Kong (HKU), Hong Kong
H
Hao Yu
Southern University of Science and Technology, Shenzhen, China
N
Ngai Wong
The University of Hong Kong (HKU), Hong Kong