Lung-SRAD: Spectral-Aware Regularized Audio DASS with Dual-Axis Patch-Mix Contrastive Learning for Respiratory Sound Classification

📅 2026-06-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing self-attention-based models for respiratory sound classification, which suffer from low-pass filtering effects that hinder effective modeling of localized abnormal acoustic features. To overcome this, the work introduces state space models (SSMs) to the task for the first time and proposes a spectrogram-aware regularization technique to better preserve mid- and high-frequency components. Additionally, a dual-axis Patch-Mix contrastive learning mechanism tailored for audio SSMs is designed to enhance feature discriminability. Evaluated on the ICBHI benchmark, the proposed method achieves a composite score of 64.48%, representing a 5% improvement over the Audio Spectrogram Transformer baseline, thereby validating the effectiveness of the proposed architecture and training strategies.
📝 Abstract
Recent respiratory sound classification (RSC) studies largely rely on CLS-token driven self-attention architectures such as the Audio Spectrogram Transformer (AST). While effective at modeling global context, recent analyses suggest a low-pass filtering behavior that may reduce sensitivity to localized abnormal patterns. In this work, we investigate State Space Models (SSMs) as an alternative backbone for RSC. Using the Distilled Audio State Space model, we analyze intermediate representations through spectral response curves and observe stronger preservation of mid-to-high spatial-frequency components. Based on these observations, we introduce spectral-aware layer regularization using Gaussian convolution applied to selected layers. We further propose Dual-Axis Patch-Mix contrastive learning tailored to SSM-based audio models for robust representation learning. Experiments on the ICBHI benchmark show that our approach achieves 64.48% score, outperforming the AST baseline by 5%. Code is available at https://github.com/RSC-Toolkit/Lung-SRAD.
Problem

Research questions and friction points this paper is trying to address.

Respiratory Sound Classification
Self-Attention
Low-Pass Filtering
Abnormal Pattern Sensitivity
Audio Spectrogram Transformer
Innovation

Methods, ideas, or system contributions that make the work stand out.

State Space Models
Spectral-Aware Regularization
Dual-Axis Patch-Mix
Contrastive Learning
Respiratory Sound Classification
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hemansh Shridhar
RSC LAB, MODULABS, Republic of Korea
M
Miika Toikkanen
RSC LAB, MODULABS, Republic of Korea
J
June-Woo Kim
Department of Electronic Engineering, Wonkwang University, Republic of Korea; AI Convergence Research Institute, Wonkwang University, Republic of Korea