Institution profile

Air University

Academic institutionasia · pk
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

Aug 06, 2026

This work addresses the limitations of existing test-time scaling methods, which often suffer from insufficient sampling diversity or reliance on external verifiers. The authors propose a verifier-free breadth–depth optimization framework that, during inference, iteratively performs self-critique and refinement across multiple reasoning trajectories, followed by majority voting to aggregate final answers. By eliminating dependence on external reward models, the approach effectively balances broad exploration with deep reasoning, significantly enhancing both computational efficiency and accuracy. Empirical results demonstrate substantial improvements over greedy decoding, conventional majority voting, and state-of-the-art verifier-based sampling strategies on mathematical reasoning benchmarks such as MATH500 and AMC; for instance, Qwen2.5-1.5B achieves an accuracy of 32.5% on AMC, up from 25.0%.

0 citationsRead paper

Enhancing Adversarial Robustness in Network Intrusion Detection: A Layer-wise Adaptive Regularization Approach

May 09, 2026

This work addresses the vulnerability of existing neural network–based intrusion detection systems to gradient-based adversarial attacks and the limitations of conventional adversarial training in interpretability and defense efficacy. The authors propose LARAR, a novel approach that integrates hierarchical vulnerability analysis, adaptive regularization, and an auxiliary classifier to quantify and optimize layer-wise vulnerability during adversarial training. LARAR introduces, for the first time, an interpretable hierarchical vulnerability score that effectively identifies critical vulnerable layers, thereby reducing computational overhead and enabling early detection of adversarial samples. Experimental results on the UNSW-NB15 dataset demonstrate that LARAR achieves a clean accuracy of 95.01% while significantly enhancing robustness against FGSM, PGD, and transfer attacks.

0 citationsRead paper

Multimodal Multi-Agent Ransomware Analysis Using AutoGen

Jan 28, 2026

Existing ransomware detection methods struggle to handle the complexity and variability of evolving ransomware families when relying solely on static, heuristic, or behavioral analysis. This work proposes the first multimodal, multi-agent collaborative detection framework that integrates static, dynamic, and network-based features. Specialized agents extract heterogeneous features, which are then fused by a dedicated fusion agent and fed into a Transformer-based classifier for ransomware family identification. The framework incorporates a confidence-aware abstention mechanism and an adaptive feedback loop to iteratively refine feature representations. Experimental results on large-scale datasets demonstrate a Macro-F1 score of 0.936, significantly reduced calibration error, an agent quality improvement exceeding 0.75, and an overall composite score of approximately 0.88—all achieved without fine-tuning any language model.

0 citationsRead paper
Recent publications

Latest Papers

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

Aug 06, 2026

This work addresses the limitations of existing test-time scaling methods, which often suffer from insufficient sampling diversity or reliance on external verifiers. The authors propose a verifier-free breadth–depth optimization framework that, during inference, iteratively performs self-critique and refinement across multiple reasoning trajectories, followed by majority voting to aggregate final answers. By eliminating dependence on external reward models, the approach effectively balances broad exploration with deep reasoning, significantly enhancing both computational efficiency and accuracy. Empirical results demonstrate substantial improvements over greedy decoding, conventional majority voting, and state-of-the-art verifier-based sampling strategies on mathematical reasoning benchmarks such as MATH500 and AMC; for instance, Qwen2.5-1.5B achieves an accuracy of 32.5% on AMC, up from 25.0%.

0 citationsRead paper

Enhancing Adversarial Robustness in Network Intrusion Detection: A Layer-wise Adaptive Regularization Approach

May 09, 2026

This work addresses the vulnerability of existing neural network–based intrusion detection systems to gradient-based adversarial attacks and the limitations of conventional adversarial training in interpretability and defense efficacy. The authors propose LARAR, a novel approach that integrates hierarchical vulnerability analysis, adaptive regularization, and an auxiliary classifier to quantify and optimize layer-wise vulnerability during adversarial training. LARAR introduces, for the first time, an interpretable hierarchical vulnerability score that effectively identifies critical vulnerable layers, thereby reducing computational overhead and enabling early detection of adversarial samples. Experimental results on the UNSW-NB15 dataset demonstrate that LARAR achieves a clean accuracy of 95.01% while significantly enhancing robustness against FGSM, PGD, and transfer attacks.

0 citationsRead paper

Multimodal Multi-Agent Ransomware Analysis Using AutoGen

Jan 28, 2026

Existing ransomware detection methods struggle to handle the complexity and variability of evolving ransomware families when relying solely on static, heuristic, or behavioral analysis. This work proposes the first multimodal, multi-agent collaborative detection framework that integrates static, dynamic, and network-based features. Specialized agents extract heterogeneous features, which are then fused by a dedicated fusion agent and fed into a Transformer-based classifier for ransomware family identification. The framework incorporates a confidence-aware abstention mechanism and an adaptive feedback loop to iteratively refine feature representations. Experimental results on large-scale datasets demonstrate a Macro-F1 score of 0.936, significantly reduced calibration error, an agent quality improvement exceeding 0.75, and an overall composite score of approximately 0.88—all achieved without fine-tuning any language model.

0 citationsRead paper