Institution profile

Pakistan Institute of Engineering and Applied Sciences

Academic institutionasia · pk
Official website
Research library12linked papers
Opportunities0open roles
Selected work

Representative Papers

Deepfake News Detection: A Multimodal Framework Integrating LipNet, DeepSpeech and ResNET for Enhanced Audio-Visual Analysis

Jul 22, 2026

This study addresses the severe threat posed by highly realistic fake news videos generated by generative AI to the authenticity of digital media. To counter this challenge, the authors propose a multimodal deepfake detection method that innovatively integrates lip motion, speech, and facial visual features. Temporal and semantic information from these modalities are extracted using LipNet, DeepSpeech2, and ResNet18—augmented with BlazeFace for face detection—respectively. A decision-level fusion is then performed via an ensemble classifier combining Random Forest, Multilayer Perceptron (MLP), and Long Short-Term Memory (LSTM) networks. Evaluated on the FakeAVCeleb dataset, the proposed approach achieves a detection accuracy of 94%, significantly outperforming existing baselines and demonstrating enhanced robustness and generalization capability.

0 citationsRead paper

MUTEX: Leveraging Multilingual Transformers and Conditional Random Fields for Enhanced Urdu Toxic Span Detection

Mar 05, 2026

This study addresses the long-standing limitation in Urdu toxic text detection, which has predominantly focused on sentence-level classification while neglecting fine-grained identification of toxic spans. To overcome challenges posed by scarce annotated data, code-mixing, and rich morphological variation, the authors present the first token-level annotated dataset for toxic span detection in Urdu. They propose a sequence labeling framework that integrates XLM-RoBERTa with a Conditional Random Field (CRF) layer to enable context-aware, fine-grained toxicity detection across multiple domains. Evaluated on social media posts, news articles, and YouTube comments, the approach achieves a token-level F1 score of 60%, establishing the first supervised baseline for Urdu toxic content analysis. The method demonstrates both strong performance and interpretability, offering a significant advancement in multilingual online safety research.

0 citationsRead paper

ULTRA:Urdu Language Transformer-based Recommendation Architecture

Feb 12, 2026

This work addresses the challenge of capturing user semantic intent in Urdu news recommendation, a low-resource language setting where existing systems struggle—particularly with queries of varying lengths. To overcome this limitation, the authors propose an adaptive semantic recommendation framework featuring a dual-channel embedding architecture and a query-length-aware dynamic routing mechanism. Short queries are processed using title-level semantic representations, while long queries leverage full-article representations, enabling fine-grained semantic alignment. The approach integrates optimized Transformer-based embeddings with tailored pooling strategies and is evaluated on a large-scale Urdu news corpus. Experimental results demonstrate that the proposed method improves recommendation accuracy by over 90% compared to single-pipeline baselines, significantly enhancing relevance and adaptability across diverse query types.

0 citationsRead paper
Recent publications

Latest Papers

Deepfake News Detection: A Multimodal Framework Integrating LipNet, DeepSpeech and ResNET for Enhanced Audio-Visual Analysis

Jul 22, 2026

This study addresses the severe threat posed by highly realistic fake news videos generated by generative AI to the authenticity of digital media. To counter this challenge, the authors propose a multimodal deepfake detection method that innovatively integrates lip motion, speech, and facial visual features. Temporal and semantic information from these modalities are extracted using LipNet, DeepSpeech2, and ResNet18—augmented with BlazeFace for face detection—respectively. A decision-level fusion is then performed via an ensemble classifier combining Random Forest, Multilayer Perceptron (MLP), and Long Short-Term Memory (LSTM) networks. Evaluated on the FakeAVCeleb dataset, the proposed approach achieves a detection accuracy of 94%, significantly outperforming existing baselines and demonstrating enhanced robustness and generalization capability.

0 citationsRead paper

MUTEX: Leveraging Multilingual Transformers and Conditional Random Fields for Enhanced Urdu Toxic Span Detection

Mar 05, 2026

This study addresses the long-standing limitation in Urdu toxic text detection, which has predominantly focused on sentence-level classification while neglecting fine-grained identification of toxic spans. To overcome challenges posed by scarce annotated data, code-mixing, and rich morphological variation, the authors present the first token-level annotated dataset for toxic span detection in Urdu. They propose a sequence labeling framework that integrates XLM-RoBERTa with a Conditional Random Field (CRF) layer to enable context-aware, fine-grained toxicity detection across multiple domains. Evaluated on social media posts, news articles, and YouTube comments, the approach achieves a token-level F1 score of 60%, establishing the first supervised baseline for Urdu toxic content analysis. The method demonstrates both strong performance and interpretability, offering a significant advancement in multilingual online safety research.

0 citationsRead paper

ULTRA:Urdu Language Transformer-based Recommendation Architecture

Feb 12, 2026

This work addresses the challenge of capturing user semantic intent in Urdu news recommendation, a low-resource language setting where existing systems struggle—particularly with queries of varying lengths. To overcome this limitation, the authors propose an adaptive semantic recommendation framework featuring a dual-channel embedding architecture and a query-length-aware dynamic routing mechanism. Short queries are processed using title-level semantic representations, while long queries leverage full-article representations, enabling fine-grained semantic alignment. The approach integrates optimized Transformer-based embeddings with tailored pooling strategies and is evaluated on a large-scale Urdu news corpus. Experimental results demonstrate that the proposed method improves recommendation accuracy by over 90% compared to single-pipeline baselines, significantly enhancing relevance and adaptability across diverse query types.

0 citationsRead paper