Institution profile

University of Nottingham Malaysia

Academic institutionasia · my
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation

Nov 23, 2025

Audio classifiers suffer from domain shift under acoustic environmental variations, yet existing test-time adaptation (TTA) studies predominantly evaluate performance under static or mismatched noise conditions, failing to model the diversity of real-world degradations. To address this, we propose DHAuDS—a novel, dynamic heterogeneous audio degradation benchmark specifically designed for audio TTA evaluation. Built upon four core datasets including UrbanSound8K-C, DHAuDS synthesizes degraded samples via dynamically modulated intensity control and multi-type noise superposition. It establishes four standardized benchmarks, introduces 14 differentiated evaluation metrics, and defines dynamic mixed-domain noise configurations. We conduct 124 reproducible experiments. As the first systematic framework for audio TTA, DHAuDS enables fair, cross-domain, and robust assessment of TTA methods under diverse, realistic audio degradations—substantially enhancing the comprehensiveness and credibility of audio model generalization evaluation.

0 citationsRead paper

AMAuT: A Flexible and Efficient Multiview Audio Transformer Framework Trained from Scratch

Oct 22, 2025

Existing foundational audio models (e.g., SSAST, EAT, HuBERT) suffer from limited generalizability and reusability due to fixed sampling rates and input duration constraints. To address this, we propose AMAuT—the first scratch-trained multi-view audio Transformer framework supporting arbitrary sampling rates and variable-length audio inputs. Our method introduces four key innovations: (1) an augmentation-driven learning paradigm; (2) a Conv1-Conv7-Conv1 bottleneck architecture for efficient spectral-temporal feature extraction; (3) a dual-token mechanism (CLS + TAL) capturing bidirectional contextual dependencies; and (4) test-time adaptive augmentation (TTA²) for robust inference. AMAuT requires no pretraining, instead leveraging multi-view data augmentation and contrastive learning to enhance robustness. Evaluated on five public benchmarks, it achieves up to 99.8% accuracy while incurring less than 3% of the training cost of comparable pretrained models. This yields substantial improvements in flexibility, cross-dataset generalization, and feasibility for edge deployment.

0 citationsRead paper

An Investigation of Test-time Adaptation for Audio Classification under Background Noise

Jul 21, 2025

To address domain shift induced by background noise in audio classification, this paper proposes CoNMix, a noise-robust test-time adaptation (TTA) method. Unlike existing test-time training (TTT) and test-time entropy minimization (TENT) approaches, CoNMix is the first TTA framework tailored to audio classification under noisy domain shifts, dynamically adapting model parameters during inference using unlabeled test samples. Evaluated on AudioMNIST and SpeechCommands under diverse noise types and signal-to-noise ratios, CoNMix consistently outperforms baseline methods, achieving a minimum error rate of 5.31% on AudioMNIST. These results demonstrate its strong generalization capability and practical efficacy. This work establishes a novel paradigm and an effective technical pathway for TTA in audio classification, advancing robustness to real-world acoustic perturbations.

0 citationsRead paper

Mobile Image Analysis Application for Mantoux Skin Test

Jun 22, 2025

Traditional tuberculin skin test (TST) interpretation relies on manual measurement of induration diameter, suffering from high subjectivity, low patient return rates, significant discomfort, and elevated misdiagnosis risk. To address these limitations, we propose a mobile-based TST image analysis system: low-cost adhesive stickers enable scale calibration; augmented reality (ARCore) spatial localization, DeepLabv3 semantic segmentation, and edge-optimization algorithms jointly achieve automatic induration detection and sub-millimeter quantitative measurement. By avoiding computationally intensive 3D reconstruction, the method balances accuracy with practical deployability. Clinical validation demonstrates a mean measurement error <0.5 mm and substantially improved inter-rater reliability versus manual assessment (ICC = 0.98 vs. 0.72). The system enhances diagnostic accuracy and accessibility, particularly enabling standardized tuberculosis screening in resource-constrained settings.

0 citationsRead paper
Recent publications

Latest Papers

DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation

Nov 23, 2025

Audio classifiers suffer from domain shift under acoustic environmental variations, yet existing test-time adaptation (TTA) studies predominantly evaluate performance under static or mismatched noise conditions, failing to model the diversity of real-world degradations. To address this, we propose DHAuDS—a novel, dynamic heterogeneous audio degradation benchmark specifically designed for audio TTA evaluation. Built upon four core datasets including UrbanSound8K-C, DHAuDS synthesizes degraded samples via dynamically modulated intensity control and multi-type noise superposition. It establishes four standardized benchmarks, introduces 14 differentiated evaluation metrics, and defines dynamic mixed-domain noise configurations. We conduct 124 reproducible experiments. As the first systematic framework for audio TTA, DHAuDS enables fair, cross-domain, and robust assessment of TTA methods under diverse, realistic audio degradations—substantially enhancing the comprehensiveness and credibility of audio model generalization evaluation.

0 citationsRead paper

AMAuT: A Flexible and Efficient Multiview Audio Transformer Framework Trained from Scratch

Oct 22, 2025

Existing foundational audio models (e.g., SSAST, EAT, HuBERT) suffer from limited generalizability and reusability due to fixed sampling rates and input duration constraints. To address this, we propose AMAuT—the first scratch-trained multi-view audio Transformer framework supporting arbitrary sampling rates and variable-length audio inputs. Our method introduces four key innovations: (1) an augmentation-driven learning paradigm; (2) a Conv1-Conv7-Conv1 bottleneck architecture for efficient spectral-temporal feature extraction; (3) a dual-token mechanism (CLS + TAL) capturing bidirectional contextual dependencies; and (4) test-time adaptive augmentation (TTA²) for robust inference. AMAuT requires no pretraining, instead leveraging multi-view data augmentation and contrastive learning to enhance robustness. Evaluated on five public benchmarks, it achieves up to 99.8% accuracy while incurring less than 3% of the training cost of comparable pretrained models. This yields substantial improvements in flexibility, cross-dataset generalization, and feasibility for edge deployment.

0 citationsRead paper

An Investigation of Test-time Adaptation for Audio Classification under Background Noise

Jul 21, 2025

To address domain shift induced by background noise in audio classification, this paper proposes CoNMix, a noise-robust test-time adaptation (TTA) method. Unlike existing test-time training (TTT) and test-time entropy minimization (TENT) approaches, CoNMix is the first TTA framework tailored to audio classification under noisy domain shifts, dynamically adapting model parameters during inference using unlabeled test samples. Evaluated on AudioMNIST and SpeechCommands under diverse noise types and signal-to-noise ratios, CoNMix consistently outperforms baseline methods, achieving a minimum error rate of 5.31% on AudioMNIST. These results demonstrate its strong generalization capability and practical efficacy. This work establishes a novel paradigm and an effective technical pathway for TTA in audio classification, advancing robustness to real-world acoustic perturbations.

0 citationsRead paper

Mobile Image Analysis Application for Mantoux Skin Test

Jun 22, 2025

Traditional tuberculin skin test (TST) interpretation relies on manual measurement of induration diameter, suffering from high subjectivity, low patient return rates, significant discomfort, and elevated misdiagnosis risk. To address these limitations, we propose a mobile-based TST image analysis system: low-cost adhesive stickers enable scale calibration; augmented reality (ARCore) spatial localization, DeepLabv3 semantic segmentation, and edge-optimization algorithms jointly achieve automatic induration detection and sub-millimeter quantitative measurement. By avoiding computationally intensive 3D reconstruction, the method balances accuracy with practical deployability. Clinical validation demonstrates a mean measurement error <0.5 mm and substantially improved inter-rater reliability versus manual assessment (ICC = 0.98 vs. 0.72). The system enhances diagnostic accuracy and accessibility, particularly enabling standardized tuberculosis screening in resource-constrained settings.

0 citationsRead paper