Institution profile

LNM Institute of Information Technology

Academic institutionasia · in
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

ECG-Lens: Benchmarking ML & DL Models on PTB-XL Dataset

Apr 17, 2026

This study addresses the application of automated electrocardiogram (ECG) classification in cardiovascular disease diagnosis by systematically evaluating the performance of various machine learning and deep learning models on real-world clinical data. Leveraging the PTB-XL dataset, the authors utilize raw 12-lead ECG signals augmented with stationary wavelet transform (SWT) and compare multiple models—including decision trees, random forests, logistic regression, a simple CNN, LSTM, and a newly proposed complex CNN architecture named ECG-Lens. Experimental results demonstrate that ECG-Lens achieves an accuracy of 80% and a ROC-AUC of 90% without requiring manual feature extraction, significantly outperforming all other evaluated methods. This work establishes ECG-Lens as a new high-performance benchmark for end-to-end automated ECG classification.

0 citationsRead paper

Reconstruction Guided Few-shot Network For Remote Sensing Image Classification

Jan 12, 2026

This work addresses the challenges of few-shot classification in remote sensing imagery, where labeled data are scarce and land cover exhibits high variability. To this end, the authors propose a Reconstruction-Guided Few-Shot Network (RGFS-Net), which integrates masked image reconstruction as an auxiliary task within a few-shot learning framework to encourage the model to learn semantically rich and spatially consistent feature representations. Notably, RGFS-Net requires no modification to the backbone architecture and is compatible with standard convolutional networks. By enforcing feature consistency constraints, the method enhances both retention of knowledge from seen classes and generalization to unseen classes. Experimental results on the EuroSAT and PatternNet datasets demonstrate that RGFS-Net consistently outperforms existing baselines under both 1-shot and 5-shot settings.

0 citationsRead paper

Two-Stage Vision Transformer for Image Restoration: Colorization Pretraining + Residual Upsampling

Dec 02, 2025

To address the limited modeling capacity of Vision Transformers (ViTs) in single-image super-resolution (SISR), this paper proposes a two-stage ViT framework. In Stage I, a self-supervised pre-training scheme is introduced using image colorization as a proxy task to enhance the model’s general representation capability for texture and structural priors. In Stage II, a residual high-frequency image prediction mechanism is incorporated, coupled with residual upsampling, to simplify the super-resolution learning process. Departing from conventional supervised pre-training paradigms, this work is the first to deeply integrate colorization into a ViT-based SISR architecture. Evaluated on DIV2K, the method achieves 22.90 dB PSNR and 0.712 SSIM—significantly outperforming baseline ViT-based approaches. These results validate the effectiveness of self-supervised representation learning and residual high-frequency modeling in boosting ViT performance for SISR.

0 citationsRead paper

Deepfake Detection of Singing Voices With Whisper Encodings

Jan 31, 2025

The proliferation of synthetic singing voice deepfakes in the music industry poses significant challenges for authenticating vocal content. Method: This paper proposes a deepfake detection method leveraging noise-variant features extracted from Whisper encoders. Departing from conventional approaches that exploit Whisper’s robustness, we first identify and harness its sensitivity to noise—specifically, forged singing voices induce distinctive, scale-dependent (tiny/base/small/medium) encoding variations across Whisper models. These variations are formalized as discriminative features. We further integrate CNN and ResNet34 architectures to jointly model both dry (unmixed) and mixed audio scenarios. Results: Extensive experiments demonstrate that our method achieves significantly lower equal error rates (EER) compared to state-of-the-art baselines, validating the effectiveness and generalizability of noise-variant encoding features for singing voice deepfake detection.

0 citationsRead paper
Recent publications

Latest Papers

ECG-Lens: Benchmarking ML & DL Models on PTB-XL Dataset

Apr 17, 2026

This study addresses the application of automated electrocardiogram (ECG) classification in cardiovascular disease diagnosis by systematically evaluating the performance of various machine learning and deep learning models on real-world clinical data. Leveraging the PTB-XL dataset, the authors utilize raw 12-lead ECG signals augmented with stationary wavelet transform (SWT) and compare multiple models—including decision trees, random forests, logistic regression, a simple CNN, LSTM, and a newly proposed complex CNN architecture named ECG-Lens. Experimental results demonstrate that ECG-Lens achieves an accuracy of 80% and a ROC-AUC of 90% without requiring manual feature extraction, significantly outperforming all other evaluated methods. This work establishes ECG-Lens as a new high-performance benchmark for end-to-end automated ECG classification.

0 citationsRead paper

Reconstruction Guided Few-shot Network For Remote Sensing Image Classification

Jan 12, 2026

This work addresses the challenges of few-shot classification in remote sensing imagery, where labeled data are scarce and land cover exhibits high variability. To this end, the authors propose a Reconstruction-Guided Few-Shot Network (RGFS-Net), which integrates masked image reconstruction as an auxiliary task within a few-shot learning framework to encourage the model to learn semantically rich and spatially consistent feature representations. Notably, RGFS-Net requires no modification to the backbone architecture and is compatible with standard convolutional networks. By enforcing feature consistency constraints, the method enhances both retention of knowledge from seen classes and generalization to unseen classes. Experimental results on the EuroSAT and PatternNet datasets demonstrate that RGFS-Net consistently outperforms existing baselines under both 1-shot and 5-shot settings.

0 citationsRead paper

Two-Stage Vision Transformer for Image Restoration: Colorization Pretraining + Residual Upsampling

Dec 02, 2025

To address the limited modeling capacity of Vision Transformers (ViTs) in single-image super-resolution (SISR), this paper proposes a two-stage ViT framework. In Stage I, a self-supervised pre-training scheme is introduced using image colorization as a proxy task to enhance the model’s general representation capability for texture and structural priors. In Stage II, a residual high-frequency image prediction mechanism is incorporated, coupled with residual upsampling, to simplify the super-resolution learning process. Departing from conventional supervised pre-training paradigms, this work is the first to deeply integrate colorization into a ViT-based SISR architecture. Evaluated on DIV2K, the method achieves 22.90 dB PSNR and 0.712 SSIM—significantly outperforming baseline ViT-based approaches. These results validate the effectiveness of self-supervised representation learning and residual high-frequency modeling in boosting ViT performance for SISR.

0 citationsRead paper

Deepfake Detection of Singing Voices With Whisper Encodings

Jan 31, 2025

The proliferation of synthetic singing voice deepfakes in the music industry poses significant challenges for authenticating vocal content. Method: This paper proposes a deepfake detection method leveraging noise-variant features extracted from Whisper encoders. Departing from conventional approaches that exploit Whisper’s robustness, we first identify and harness its sensitivity to noise—specifically, forged singing voices induce distinctive, scale-dependent (tiny/base/small/medium) encoding variations across Whisper models. These variations are formalized as discriminative features. We further integrate CNN and ResNet34 architectures to jointly model both dry (unmixed) and mixed audio scenarios. Results: Extensive experiments demonstrate that our method achieves significantly lower equal error rates (EER) compared to state-of-the-art baselines, validating the effectiveness and generalizability of noise-variant encoding features for singing voice deepfake detection.

0 citationsRead paper