Institution profile

University of Kurdistan

Academic institutionasia · iq
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

From Dialect Gaps to Identity Maps: Tackling Variability in Speaker Verification

Apr 21, 2025arXiv.org

This study addresses the significant performance degradation of speaker verification in multidilectal Kurdish (Kurmanji, Sorani, Hawrami). We propose a dual-path framework integrating dialect-specific modeling and cross-dialect joint training. Methodologically, we construct the first annotated speech corpus covering all three major Kurdish dialects; design dialect-aware data augmentation, adversarial dialect-invariant feature learning, and multi-task loss optimization to robustly disentangle speaker identity representations from dialectal variation. Evaluated on cross-dialect test sets using an x-vector–based speaker embedding model, our approach achieves a 38.7% relative reduction in equal error rate (EER) compared to both single-dialect baselines and general multilingual speaker verification systems. This work constitutes the first systematic solution to cross-dialect speaker verification in Kurdish and establishes a novel paradigm for low-resource, multidilectal voice biometrics.

2 citationsRead paper

Alcmean's: Unsupervised community detection using local Laplacian, automatic detection of the number of centers

Jun 08, 2026

This study addresses key limitations in complex network community detection—namely, the need to predefine the number of communities, inaccurate selection of cluster centers, and poor scalability—by proposing an unsupervised method that requires no prior specification of community count. The approach automatically identifies structurally pivotal nodes as cluster centers using local Laplacian centrality and integrates DeepWalk node embeddings with unsupervised clustering to enhance the accuracy and stability of community assignment. Extensive evaluations on multiple benchmark datasets demonstrate that the proposed method outperforms state-of-the-art algorithms by 10%–20% in both Normalized Mutual Information (NMI) and Adjusted Rand Index (ARI), while also achieving significantly higher Modularity and F1 scores, thereby confirming its effectiveness and scalability.

0 citationsRead paper

ECG-NAT: A Self-supervised Neighborhood Attention Transformer for Multi-lead Electrocardiogram Classification

May 13, 2026

This work addresses the challenges in electrocardiogram (ECG) arrhythmia classification—namely, high signal variability, strong noise interference, scarce labeled data, and the trade-off between accuracy and efficiency—by proposing ECG-NAT, a novel two-stage self-supervised learning framework. The method first employs a masked autoencoder for generative pretraining on multi-source ECG data, followed by discriminative fine-tuning that jointly optimizes supervised contrastive and cross-entropy losses. A key innovation is the introduction of a hierarchical neighborhood attention mechanism, which efficiently captures multiscale temporal features ranging from individual heartbeat morphology to global rhythm patterns. Experimental results demonstrate that ECG-NAT achieves 88.1% accuracy on standard benchmarks using only 1% of labeled data, maintaining superior classification performance while significantly reducing computational overhead, thereby making it well-suited for real-time ECG diagnostics.

0 citationsRead paper

BiCLIP: Bidirectional and Consistent Language-Image Processing for Robust Medical Image Segmentation

Feb 25, 2026

This work addresses the challenge of robust medical image segmentation under realistic clinical conditions characterized by scarce annotations and image degradations such as motion blur and low-dose CT noise. To this end, the authors propose a bidirectional vision–language fusion framework that iteratively refines multimodal interactions between image features and textual descriptions. Enhanced consistency regularization is introduced to stabilize training dynamics. Leveraging a contrastive language–image pretraining architecture, the model achieves state-of-the-art segmentation accuracy and robustness on the QaTa-COV19 and MosMedData+ benchmarks, significantly outperforming existing methods even when trained with only 1% labeled data.

0 citationsRead paper

Automatic Text Summarization (ATS) for Research Documents in Sorani Kurdish

Apr 20, 2025

This work addresses the lack of research resources for Sorani Kurdish—a low-resource language—by introducing the first academic paper summarization dataset specifically designed for scholarly literature, comprising 231 annotated papers. To establish a reproducible baseline, we propose a lightweight unsupervised model combining TF-IDF with sentence weighting. Evaluation employs both automated metrics (ROUGE-1, ROUGE-2, ROUGE-L) and rigorous human assessment by six domain-expert annotators. The best-performing configuration achieves a ROUGE-L score of 19.58%, demonstrating the feasibility of automatic summarization for Sorani Kurdish academic texts. Crucially, we release the full dataset, preprocessing scripts, and model implementations under an open-source license. This contribution provides the foundational infrastructure and a standardized, reproducible benchmark to catalyze future NLP research on Sorani Kurdish, particularly in scientific text processing and summarization.

0 citationsRead paper
Recent publications

Latest Papers

Alcmean's: Unsupervised community detection using local Laplacian, automatic detection of the number of centers

Jun 08, 2026

This study addresses key limitations in complex network community detection—namely, the need to predefine the number of communities, inaccurate selection of cluster centers, and poor scalability—by proposing an unsupervised method that requires no prior specification of community count. The approach automatically identifies structurally pivotal nodes as cluster centers using local Laplacian centrality and integrates DeepWalk node embeddings with unsupervised clustering to enhance the accuracy and stability of community assignment. Extensive evaluations on multiple benchmark datasets demonstrate that the proposed method outperforms state-of-the-art algorithms by 10%–20% in both Normalized Mutual Information (NMI) and Adjusted Rand Index (ARI), while also achieving significantly higher Modularity and F1 scores, thereby confirming its effectiveness and scalability.

0 citationsRead paper

ECG-NAT: A Self-supervised Neighborhood Attention Transformer for Multi-lead Electrocardiogram Classification

May 13, 2026

This work addresses the challenges in electrocardiogram (ECG) arrhythmia classification—namely, high signal variability, strong noise interference, scarce labeled data, and the trade-off between accuracy and efficiency—by proposing ECG-NAT, a novel two-stage self-supervised learning framework. The method first employs a masked autoencoder for generative pretraining on multi-source ECG data, followed by discriminative fine-tuning that jointly optimizes supervised contrastive and cross-entropy losses. A key innovation is the introduction of a hierarchical neighborhood attention mechanism, which efficiently captures multiscale temporal features ranging from individual heartbeat morphology to global rhythm patterns. Experimental results demonstrate that ECG-NAT achieves 88.1% accuracy on standard benchmarks using only 1% of labeled data, maintaining superior classification performance while significantly reducing computational overhead, thereby making it well-suited for real-time ECG diagnostics.

0 citationsRead paper

BiCLIP: Bidirectional and Consistent Language-Image Processing for Robust Medical Image Segmentation

Feb 25, 2026

This work addresses the challenge of robust medical image segmentation under realistic clinical conditions characterized by scarce annotations and image degradations such as motion blur and low-dose CT noise. To this end, the authors propose a bidirectional vision–language fusion framework that iteratively refines multimodal interactions between image features and textual descriptions. Enhanced consistency regularization is introduced to stabilize training dynamics. Leveraging a contrastive language–image pretraining architecture, the model achieves state-of-the-art segmentation accuracy and robustness on the QaTa-COV19 and MosMedData+ benchmarks, significantly outperforming existing methods even when trained with only 1% labeled data.

0 citationsRead paper

From Dialect Gaps to Identity Maps: Tackling Variability in Speaker Verification

Apr 21, 2025arXiv.org

This study addresses the significant performance degradation of speaker verification in multidilectal Kurdish (Kurmanji, Sorani, Hawrami). We propose a dual-path framework integrating dialect-specific modeling and cross-dialect joint training. Methodologically, we construct the first annotated speech corpus covering all three major Kurdish dialects; design dialect-aware data augmentation, adversarial dialect-invariant feature learning, and multi-task loss optimization to robustly disentangle speaker identity representations from dialectal variation. Evaluated on cross-dialect test sets using an x-vector–based speaker embedding model, our approach achieves a 38.7% relative reduction in equal error rate (EER) compared to both single-dialect baselines and general multilingual speaker verification systems. This work constitutes the first systematic solution to cross-dialect speaker verification in Kurdish and establishes a novel paradigm for low-resource, multidilectal voice biometrics.

2 citationsRead paper

Automatic Text Summarization (ATS) for Research Documents in Sorani Kurdish

Apr 20, 2025

This work addresses the lack of research resources for Sorani Kurdish—a low-resource language—by introducing the first academic paper summarization dataset specifically designed for scholarly literature, comprising 231 annotated papers. To establish a reproducible baseline, we propose a lightweight unsupervised model combining TF-IDF with sentence weighting. Evaluation employs both automated metrics (ROUGE-1, ROUGE-2, ROUGE-L) and rigorous human assessment by six domain-expert annotators. The best-performing configuration achieves a ROUGE-L score of 19.58%, demonstrating the feasibility of automatic summarization for Sorani Kurdish academic texts. Crucially, we release the full dataset, preprocessing scripts, and model implementations under an open-source license. This contribution provides the foundational infrastructure and a standardized, reproducible benchmark to catalyze future NLP research on Sorani Kurdish, particularly in scientific text processing and summarization.

0 citationsRead paper