From 80x to 385x: A Best-Matching-Unit Search at the L2 Roof, Measured Against a Symmetrically Tuned Baseline

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
通过调整SparseBin和cuSPARSE算法的四个参数,优化了自组织映射训练中的最佳匹配单元搜索,显著提高了处理速度。
📝 Abstract
Comparisons between GPU implementations are usually asymmetric: one side is tuned by its author, the other is run as found. I report a programme that tuned both a novel SOM algorithm (SparseBin) and the baseline algorithm it was being compared to (cuSPARSE). The best-matching-unit search that dominates self-organizing map training was tuned through four levers - tile size, tile-membership clustering, neuron-axis chunking and vectorised loads - reaching 5.6-10.1x per epoch over the previously published configuration at map sizes from 32x32 to 512x512, and lifting the margin over the CUDA implementation behind our earlier MEDLINE atlases from ~80x to ~385x. cuSPARSE, the implementation SparseBin is compared against, received every lever with an analogue on its side, and became 2-3x faster in the process. The tuned kernel pressed the L2 bandwidth roof at 77% of peak with every other unit at 40-65%, bounding any further lever at ~1.3x - a terminal result rather than a waypoint, and every untested lever was either capped by that bound by construction or measured null.
Problem

Research questions and friction points this paper is trying to address.

asymmetric comparisons
GPU implementations
tuning
SOM algorithm
baseline
Innovation

Methods, ideas, or system contributions that make the work stand out.

Best-Matching-Unit Search
SparseBin
L2 Bandwidth Optimization
cuSPARSE
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
A
Andrew J. Amos
College of Medicine and Dentistry, James Cook University, Townsville, Queensland, Australia