A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决现代AI工作负载中张量计算的调度瓶颈,提出FIBER架构,通过解耦线程与寄存器所有权实现动态并行和细粒度调度。
📝 Abstract
Modern GPUs increasingly integrate Tensor Cores into the execution pipeline. Although aggregate tensor throughput continues to grow, aided by an operand supply that has evolved from register-based in Ampere to redundancy-free, memory-based in Hopper and Blackwell, efficiently orchestrating the complete tensor compute pipeline for the modern AI workloads remains challenging. We identify the fundamental bottlenecks as fixed parallelism and coarse-grained scheduling, both of which are exposed by modern AI workloads that interleave diverse non-GEMM operations with GEMM. To orchestrate tensor computation efficiently, we propose FIBER, a new architecture that extends the GPU SIMT (single instruction, multiple thread) model. Its basic execution instance, the \emph{fiber}, is decoupled from private register ownership, carrying only minimal control state while accessing an SM's registers through a shared view. This enables dynamic parallelism scaling, fine-grained register-level dataflow scheduling, and offers a redundancy-free alternative for matrix operand supply. We extend the ISA, microarchitecture, and compiler to realize shared-register addressing, conflict-free operand delivery, and fiber-based program mapping. Under a typical mixed-precision LLM serving scenario, FIBER achieves a 2.25x end-to-end speedup on Ampere (1.15x for the original FP16 computation), with 1.8x and 2.09x on Hopper and Blackwell respectively, and kernel-level gains up to 2.49x.
Problem

Research questions and friction points this paper is trying to address.

Tensor Cores
Fixed Parallelism
Coarse-Grained Scheduling
GEMM Operations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Thread-Register Decoupled
Dynamic Parallelism Scaling
Fine-grained Scheduling
Redundancy-free Operand Supply
🔎 Similar Papers
No similar papers found.