Rethinking Attention Locality in Spiking Transformers

๐Ÿ“… 2026-08-09
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the lack of efficient and consistent spatially local attention mechanisms in existing Spiking Transformers. The authors propose SCLA-BCP, a novel approach that explicitly distinguishes computational locality from spatial locality for the first time. It introduces spatially contiguous local attention within non-overlapping neighboring regions, coupled with a lightweight convolutional boundary pathway, and incorporates an architecture-adaptive hierarchical locality deployment strategy. Evaluated across seven static and neuromorphic datasets, the method achieves substantial performance gains with minimal parameter and energy overheadโ€”most notably, a 9.50% absolute improvement in mAP@50 on COCO 2017 and a 3.42% gain in mIoU on ADE20K.
๐Ÿ“ Abstract
Spiking Transformers provide a promising paradigm for efficient visual processing with spike-driven computation, yet their Softmax-free Spiking Self-Attention (SSA) struggles to establish spatially localized token interactions. Although existing locality-enhanced SSA methods improve accuracy, it remains unclear whether they consistently induce spatial locality across layers and different Spiking Transformer architectures. Through Mean Attention Distance (MAD) analysis, we reveal that computational locality does not necessarily translate into spatial locality and show that uniformly applying the same locality enhancement overlooks architecture-dependent deployment requirements. Motivated by these observations, we propose Spatially Contiguous Local Attention with Boundary Continuity Pathway (SCLA-BCP). SCLA computes attention within non-overlapping regions of spatially adjacent tokens, while BCP facilitates cross-boundary information exchange through a lightweight convolutional pathway. Furthermore, we develop a hierarchical locality deployment strategy to effectively apply SCLA-BCP across the two major Spiking Transformer architectures. Extensive experiments on seven static and neuromorphic datasets covering classification, detection, and segmentation demonstrate consistent improvements with limited parameter and energy overhead. Notably, our approach improves mAP@50 by up to 9.50% on COCO 2017 and mIoU by up to 3.42% on ADE20K. Visualizations, MAD analysis, and ablation studies further validate its effectiveness.
Problem

Research questions and friction points this paper is trying to address.

Spiking Transformers
Spatial Locality
Self-Attention
Attention Locality
Mean Attention Distance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spiking Transformers
Spatial Locality
Local Attention
Boundary Continuity Pathway
Mean Attention Distance
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Z
Zeqi Zheng
Zhejiang University
Z
Zizheng Zhu
Zhejiang University
Y
Yuping Yan
Westlake University
Wenxuan Pan
Wenxuan Pan
Institute of Automation, Chinese Academy of Sciences
Brain-inspired Intelligence
Zhaofei Yu
Zhaofei Yu
Peking University
Brain-inspired ComputingSpiking Neural NetworksComputational Neuroscience
Y
Yaochu Jin
Westlake University