AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出AsmEvo,通过在AMD GPU汇编级进行优化并验证功能等价性,解决已编译代码对象的进一步优化问题。
📝 Abstract
High-performance ML systems increasingly rely on GPU kernels whose editable source is unavailable, generated, or too distant from final machine code to expose remaining optimizations. Existing LLM kernel optimizers and autotuners mainly operate on CUDA, Triton, HIP, or tensor-program source and validate against reference implementations. We study a stricter setting: optimizing an already compiled AMDGPU code object, where the deployed binary is the only behavioral oracle. We present AsmEvo, an agentic assembly-level optimizer for AMD GPU kernels. Given an AMDGPU code object K0, AsmEvo reconstructs a reassemblable representation, proposes low-level edits with a long-horizon agent, rebuilds an ABI-preserving optimized object, and accepts candidates only after differential verification against K0 under identical launches. AsmEvo combines code-object recovery, metadata-aware rebuilding, profiling-guided hot-window editing, correctness-gated timing, and conservative in-place patch fallback. We conduct extensive experiments with AsmEvo on various AMD GPU kernels. On MI308X, AsmEvo improves 29 of 30 selected KernelBench kernels, reaching 1.35x geometric-mean and 3.88x maximum speedup. On MI300X production workloads, it improves all evaluated AITer binaries and vLLM/SGLang Triton assembly kernels, reaching 1.09x/1.31x and 1.18x/1.34x geometric-mean/maximum speedups, respectively, while preserving functional equivalence.
Problem

Research questions and friction points this paper is trying to address.

AMD GPU Kernels
Assembly-Level Optimization
Functional Equivalence Verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

assembly-level optimization
functional equivalence verification
AMD GPU kernels
code-object recovery
metadata-aware rebuilding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ji Liu
Ji Liu
AMD
Autonomous Driving/PerceptionLLM/LMMAIGCCompute VisionModel Optimization.
Puyuan Yang
Puyuan Yang
University of Science and Technology of China, USTC; Alibaba AliCloud
Storage SystemDatabaseFlashHybrid StorageCloud Storage
Rongzhang Zheng
Rongzhang Zheng
Advanced Micro Devices, Inc. (AMD)
F
Fan Wang
Advanced Micro Devices, Inc. (AMD)
J
Jinglin Wang
Advanced Micro Devices, Inc. (AMD), Southern University of Science and Technology, Shenzhen, China
M
Muhammad A. Awad
Advanced Micro Devices, Inc. (AMD)
M
Mortis Huang
Advanced Micro Devices, Inc. (AMD)
A
Andy Chang
Advanced Micro Devices, Inc. (AMD)
Zekai Li
Zekai Li
PhD Student at UC San Diego
efficient deep learningdata-centric AI
Zeping Li
Zeping Li
Phd student in Fudan University, Financial Technology Group
LLM and KG
Z
Zihao An
Advanced Micro Devices, Inc. (AMD)
Y
Yue Liu
Advanced Micro Devices, Inc. (AMD)
Y
Yuchen Yang
Advanced Micro Devices, Inc. (AMD)
J
Jianghui Wang
Advanced Micro Devices, Inc. (AMD)
C
Chushi Chen
Advanced Micro Devices, Inc. (AMD)
Ziqiong Liu
Ziqiong Liu
MIND
Multimedia SearchComputer VisionMachine Learning
F
Fuwei Yang
Advanced Micro Devices, Inc. (AMD)
D
Dong Li
Advanced Micro Devices, Inc. (AMD)
W
Wen Heng Chung
Advanced Micro Devices, Inc. (AMD)
Shengcai Liu
Shengcai Liu
Southern University of Science and Technology
Learn to OptimizeLLM+Optimization
Emad Barsoum
Emad Barsoum
AMD, Columbia University
Generative AIFoundation ModelsAgentic AIComputer VisionML Frameworks