KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出KernelArc,一种多代理框架,通过并行运行的策略特化代理和特定协调机制来解决GPU内核在异构工作负载下的优化问题。
📝 Abstract
We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-agent state with plateau-triggered drafting. We evaluate \kernelarc{} on NVIDIA H100 and B200 GPUs using category-representative SOL-ExecBench workloads. The resulting implementations span custom BF16 GEMM, static cuBLASLt Expert-API configuration tables, fused mixture-of-experts backward, shape-gated decoder-layer fusion, native NVFP4 grouped-query attention, and paged prefill attention. At the public SOL-ExecBench leaderboard snapshot recorded on July~30, 2026, these submissions ranked first on representative L1, L2, Quantization, and FlashInfer tasks. The trajectories support the paper's central motivation: shared multi-agent search can broaden exploration and reach stronger incumbents within a fixed candidate budget, while the value of individual coordination features depends on the kernel and optimization stage.
Problem

Research questions and friction points this paper is trying to address.

GPU kernel optimization
heterogeneous workloads
multi-agent framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-agent framework
GPU kernel optimization
heterogeneous workloads
conclusions-only shared memory
deterministic benchmark guard