Institution profile

Institute of Engineering

Academic institution
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Character Recognition of Nepali Number Plate

Jun 27, 2026

This study addresses the challenge of automatic license plate recognition in Nepal, where Devanagari script characters pose significant difficulties. The work proposes the first end-to-end license plate recognition system tailored to local, complex real-world scenarios. It integrates YOLO-based models for detecting both license plate regions and individual character locations, coupled with a dedicated CNN classifier trained specifically to recognize 34 Devanagari characters. Through extensive data augmentation and targeted training on embossed plates, the system substantially enhances generalization under diverse real-world conditions. Evaluated on a realistic dataset encompassing variations in lighting, font styles, and plate structures, the approach achieves a character-level recognition accuracy of up to 93%, offering an efficient and scalable solution for intelligent traffic management in Nepal.

0 citationsRead paper

Neural Speaker Diarization via Multilingual Training: Evaluation on Low-Resource Nepali-Hindi Speech

Jun 21, 2026

This work addresses the significant performance degradation in speaker diarization for low-resource languages such as Nepali and Hindi, primarily due to scarce annotated data. The study introduces DiaPer, the first application of the Perceiver architecture to this task, which leverages an attractor-based mechanism and is evaluated against EEND-EDA within a multilingual joint training framework. End-to-end training is conducted using LibriSpeech, VoxCeleb, and a newly collected NeHi corpus. On the NeHi test set, DiaPer achieves diarization error rates (DER) of 3.28%, 2.02%, 4.05%, and 4.76% for scenarios involving two, three, four speakers, and mixed conditions, respectively—substantially outperforming baseline models. The results demonstrate that DiaPer effectively mitigates language bias and enhances generalization in low-resource multilingual settings.

0 citationsRead paper

EEG-based AI-BCI Wheelchair Advancement: Hybrid Deep Learning with Motor Imagery for Brain Computer Interface

Sep 29, 2025

To address insufficient EEG signal decoding accuracy in motor imagery (MI)-based brain–computer interface (BCI) wheelchair control, this study proposes an attention-enhanced hybrid deep learning model integrating bidirectional LSTM and bidirectional GRU (BiLSTM-BiGRU). The architecture effectively captures long-range temporal dependencies in EEG sequences while adaptively amplifying discriminative feature weights via a learnable attention mechanism. Evaluated on standard MI-EEG datasets under rigorous five-fold cross-validation, the model achieves 92.26% single-trial classification accuracy and 90.13% mean cross-validated accuracy—outperforming XGBoost, EEGNet, and Transformer baselines. Furthermore, a real-time visualization simulation interface was implemented using Tkinter, enabling intuitive wheelchair navigation via left/right hand MI commands. This work demonstrates the efficacy and practicality of attention-augmented dual-gated recurrent architectures for lightweight, deployable BCI decoding, offering a novel pathway toward real-world intelligent assistive systems.

0 citationsRead paper

Lightweight MobileNetV1+GRU for ECG Biometric Authentication: Federated and Adversarial Evaluation

Sep 21, 2025

To address the challenges of poor real-time performance, privacy leakage, and vulnerability to adversarial attacks in electrocardiogram (ECG)-based biometric authentication for wearable devices, this paper proposes a lightweight secure authentication framework. The method integrates MobileNetV1 and GRU for low-latency time-frequency feature extraction, employs federated learning to ensure privacy-preserving distributed training, and incorporates robust preprocessing with 20-dB Gaussian noise alongside FGSM-based adversarial evaluation. Evaluated on four benchmark datasets—ECGID, MIT-BIH, CYBHi, and PTB—the framework achieves 99.34% accuracy, F1-score of 0.9923, equal error rate (EER) of 0.00013, and ROC-AUC of 0.9999. However, under FGSM attacks, accuracy drops to 80%, exposing a key robustness limitation. This work establishes a reproducible technical pathway and empirical benchmark for privacy-enhanced, spoof-resistant ECG authentication deployable at the edge.

0 citationsRead paper

NepaliGPT: A Generative Language Model for the Nepali Language

Jun 19, 2025

Nepali lacks dedicated generative large language models (LLMs), severely constraining downstream task research. Method: We introduce NepaliGPT—the first open-source, autoregressive LLM specifically designed for Nepali. Our approach comprises (i) constructing a large-scale, domain-diverse Devanagari-script text corpus; (ii) designing the first Nepali question-answering benchmark (4,296 QA pairs); and (iii) developing a customized evaluation framework incorporating ROUGE scores, causal coherence, and causal consistency metrics. Results: NepaliGPT achieves a perplexity of 26.32 on text generation, ROUGE-1 of 0.2604, causal coherence of 81.25%, and causal consistency of 85.41%. This work fills a critical gap in foundational Nepali LLMs and establishes a reusable paradigm—including data, model architecture, and evaluation methodology—for low-resource language LLM development.

0 citationsRead paper
Recent publications

Latest Papers

Character Recognition of Nepali Number Plate

Jun 27, 2026

This study addresses the challenge of automatic license plate recognition in Nepal, where Devanagari script characters pose significant difficulties. The work proposes the first end-to-end license plate recognition system tailored to local, complex real-world scenarios. It integrates YOLO-based models for detecting both license plate regions and individual character locations, coupled with a dedicated CNN classifier trained specifically to recognize 34 Devanagari characters. Through extensive data augmentation and targeted training on embossed plates, the system substantially enhances generalization under diverse real-world conditions. Evaluated on a realistic dataset encompassing variations in lighting, font styles, and plate structures, the approach achieves a character-level recognition accuracy of up to 93%, offering an efficient and scalable solution for intelligent traffic management in Nepal.

0 citationsRead paper

Neural Speaker Diarization via Multilingual Training: Evaluation on Low-Resource Nepali-Hindi Speech

Jun 21, 2026

This work addresses the significant performance degradation in speaker diarization for low-resource languages such as Nepali and Hindi, primarily due to scarce annotated data. The study introduces DiaPer, the first application of the Perceiver architecture to this task, which leverages an attractor-based mechanism and is evaluated against EEND-EDA within a multilingual joint training framework. End-to-end training is conducted using LibriSpeech, VoxCeleb, and a newly collected NeHi corpus. On the NeHi test set, DiaPer achieves diarization error rates (DER) of 3.28%, 2.02%, 4.05%, and 4.76% for scenarios involving two, three, four speakers, and mixed conditions, respectively—substantially outperforming baseline models. The results demonstrate that DiaPer effectively mitigates language bias and enhances generalization in low-resource multilingual settings.

0 citationsRead paper

EEG-based AI-BCI Wheelchair Advancement: Hybrid Deep Learning with Motor Imagery for Brain Computer Interface

Sep 29, 2025

To address insufficient EEG signal decoding accuracy in motor imagery (MI)-based brain–computer interface (BCI) wheelchair control, this study proposes an attention-enhanced hybrid deep learning model integrating bidirectional LSTM and bidirectional GRU (BiLSTM-BiGRU). The architecture effectively captures long-range temporal dependencies in EEG sequences while adaptively amplifying discriminative feature weights via a learnable attention mechanism. Evaluated on standard MI-EEG datasets under rigorous five-fold cross-validation, the model achieves 92.26% single-trial classification accuracy and 90.13% mean cross-validated accuracy—outperforming XGBoost, EEGNet, and Transformer baselines. Furthermore, a real-time visualization simulation interface was implemented using Tkinter, enabling intuitive wheelchair navigation via left/right hand MI commands. This work demonstrates the efficacy and practicality of attention-augmented dual-gated recurrent architectures for lightweight, deployable BCI decoding, offering a novel pathway toward real-world intelligent assistive systems.

0 citationsRead paper

Lightweight MobileNetV1+GRU for ECG Biometric Authentication: Federated and Adversarial Evaluation

Sep 21, 2025

To address the challenges of poor real-time performance, privacy leakage, and vulnerability to adversarial attacks in electrocardiogram (ECG)-based biometric authentication for wearable devices, this paper proposes a lightweight secure authentication framework. The method integrates MobileNetV1 and GRU for low-latency time-frequency feature extraction, employs federated learning to ensure privacy-preserving distributed training, and incorporates robust preprocessing with 20-dB Gaussian noise alongside FGSM-based adversarial evaluation. Evaluated on four benchmark datasets—ECGID, MIT-BIH, CYBHi, and PTB—the framework achieves 99.34% accuracy, F1-score of 0.9923, equal error rate (EER) of 0.00013, and ROC-AUC of 0.9999. However, under FGSM attacks, accuracy drops to 80%, exposing a key robustness limitation. This work establishes a reproducible technical pathway and empirical benchmark for privacy-enhanced, spoof-resistant ECG authentication deployable at the edge.

0 citationsRead paper

NepaliGPT: A Generative Language Model for the Nepali Language

Jun 19, 2025

Nepali lacks dedicated generative large language models (LLMs), severely constraining downstream task research. Method: We introduce NepaliGPT—the first open-source, autoregressive LLM specifically designed for Nepali. Our approach comprises (i) constructing a large-scale, domain-diverse Devanagari-script text corpus; (ii) designing the first Nepali question-answering benchmark (4,296 QA pairs); and (iii) developing a customized evaluation framework incorporating ROUGE scores, causal coherence, and causal consistency metrics. Results: NepaliGPT achieves a perplexity of 26.32 on text generation, ROUGE-1 of 0.2604, causal coherence of 81.25%, and causal consistency of 85.41%. This work fills a critical gap in foundational Nepali LLMs and establishes a reusable paradigm—including data, model architecture, and evaluation methodology—for low-resource language LLM development.

0 citationsRead paper