Institution profile

Information Technology University

Academic institutionasia · pk
Official website
Research library32linked papers
Opportunities0open roles
Selected work

Representative Papers

Who is Responsible? The Data, Models, Users or Regulations? Responsible Generative AI for a Sustainable Future

Jan 15, 2025

This paper addresses the implementation gap in ethical governance of generative AI (Gen AI) in the post-ChatGPT era. Methodologically, it introduces the first end-to-end Responsible Gen AI (RAI) practice framework—spanning governance, technology, evaluation, and deployment—integrating philosophical responsibility theory, eXplainable AI (XAI), benchmark alignment, cross-sector application modeling, and KPI-based quantitative assessment. It pioneers an AI-readiness-oriented testbed evaluation methodology and establishes a comprehensive RAI Key Performance Indicator (KPI) system. Contributions include: (1) systematically bridging the chasm between normative ethical principles and engineering practice while redefining accountability structures; and (2) releasing an open-source resource repository—including standards, tools, and benchmark datasets—to provide researchers, policymakers, and industry practitioners with scalable, reusable, and trustworthy implementation guidance.

1 citationsRead paper

Consensus-gated Multi-Agent Neural Architecture Search for Seismic Fault Segmentation

Aug 13, 2026

This study addresses the challenges of architectural design and high computational costs associated with traditional Neural Architecture Search (NAS) in few-shot seismic fault segmentation. We propose a Multi-Agent Consensus-Gated NAS system that pioneers code-level architecture search via Large Language Model debates, thereby eliminating predefined operation space constraints. Through multi-agent collaboration and automated verification loops, this approach enables efficient model discovery. Experiments demonstrate that training only eight candidate models—requiring approximately one GPU-day—yields an optimal 425K-parameter network achieving an F1-score of 0.578. This performance significantly surpasses larger architectures such as U-Net, validating the method’s capability to automatically construct high-performance, domain-specific networks under strict computational constraints.

0 citationsRead paper

Automated Borehole Core Analysis with Report-Derived Weak Labels and Supervised Crack Segmentation

Aug 12, 2026

This study addresses the challenge posed by the absence of pixel-level fracture annotations in core images, which hinders automated extraction of fracture spacing and associated geological features. To overcome this limitation, the authors propose a multimodal weakly supervised learning framework that leverages digital log reports to generate weak labels for fracture spacing classification and integrates these with limited human-annotated strong labels to train a fully supervised segmentation model. The approach innovatively incorporates a learnable spatial gating mechanism, combining a DINO self-supervised encoder, PiDiNet for edge detection, and Mask R-CNN for instance segmentation, alongside rule-based modules for estimating bedding angle and lithology color. Experimental results demonstrate a fracture segmentation F1 score of 0.860 (IoU 0.754), with bedding angle and lithology color predictions achieving 75.4% and 84.7% consistency, respectively, with expert log reports.

0 citationsRead paper

Weight and Height Estimation from a Single Human Image Captured in the Wild

Jul 28, 2026

This study addresses the challenge of accurately estimating height, weight, and BMI from a single image in complex real-world scenarios, where factors such as pose variation, viewing angle, and cluttered backgrounds significantly hinder performance. To tackle this problem, the authors propose a deep neural network framework that leverages multimodal inputs—specifically RGB images, depth maps, pose affinity fields, and edge maps—and integrates both single-task and multi-task learning strategies for human body shape parameter prediction from unconstrained full-body images. A key contribution is the introduction of the first publicly available in-the-wild full-body anthropometric dataset, comprising 6,105 multi-ethnic, multi-pose images. Systematic evaluation reveals that full-body images substantially outperform upper-body or facial crops, and that backbone architectures such as VGG, DenseNet, and ResNet achieve consistently higher estimation accuracy when fused with multimodal cues.

0 citationsRead paper
Recent publications

Latest Papers

Consensus-gated Multi-Agent Neural Architecture Search for Seismic Fault Segmentation

Aug 13, 2026

This study addresses the challenges of architectural design and high computational costs associated with traditional Neural Architecture Search (NAS) in few-shot seismic fault segmentation. We propose a Multi-Agent Consensus-Gated NAS system that pioneers code-level architecture search via Large Language Model debates, thereby eliminating predefined operation space constraints. Through multi-agent collaboration and automated verification loops, this approach enables efficient model discovery. Experiments demonstrate that training only eight candidate models—requiring approximately one GPU-day—yields an optimal 425K-parameter network achieving an F1-score of 0.578. This performance significantly surpasses larger architectures such as U-Net, validating the method’s capability to automatically construct high-performance, domain-specific networks under strict computational constraints.

0 citationsRead paper

Automated Borehole Core Analysis with Report-Derived Weak Labels and Supervised Crack Segmentation

Aug 12, 2026

This study addresses the challenge posed by the absence of pixel-level fracture annotations in core images, which hinders automated extraction of fracture spacing and associated geological features. To overcome this limitation, the authors propose a multimodal weakly supervised learning framework that leverages digital log reports to generate weak labels for fracture spacing classification and integrates these with limited human-annotated strong labels to train a fully supervised segmentation model. The approach innovatively incorporates a learnable spatial gating mechanism, combining a DINO self-supervised encoder, PiDiNet for edge detection, and Mask R-CNN for instance segmentation, alongside rule-based modules for estimating bedding angle and lithology color. Experimental results demonstrate a fracture segmentation F1 score of 0.860 (IoU 0.754), with bedding angle and lithology color predictions achieving 75.4% and 84.7% consistency, respectively, with expert log reports.

0 citationsRead paper

Weight and Height Estimation from a Single Human Image Captured in the Wild

Jul 28, 2026

This study addresses the challenge of accurately estimating height, weight, and BMI from a single image in complex real-world scenarios, where factors such as pose variation, viewing angle, and cluttered backgrounds significantly hinder performance. To tackle this problem, the authors propose a deep neural network framework that leverages multimodal inputs—specifically RGB images, depth maps, pose affinity fields, and edge maps—and integrates both single-task and multi-task learning strategies for human body shape parameter prediction from unconstrained full-body images. A key contribution is the introduction of the first publicly available in-the-wild full-body anthropometric dataset, comprising 6,105 multi-ethnic, multi-pose images. Systematic evaluation reveals that full-body images substantially outperform upper-body or facial crops, and that backbone architectures such as VGG, DenseNet, and ResNet achieve consistently higher estimation accuracy when fused with multimodal cues.

0 citationsRead paper

InstructMixup: Instruction-Guided Salient Patch Editing for Robust Data Augmentation

Jul 21, 2026

This work addresses the high computational cost and semantic inconsistency inherent in existing mixup-based data augmentation methods, which often impair model generalization by blending across samples. To overcome these limitations, the authors propose a single-sample, semantically coherent strong augmentation strategy: key regions are first identified via lightweight multi-scale saliency detection, then edited using instruction-guided generative models, and finally recombined with the non-salient parts of the original image, while adaptive fractal structures are injected to enhance representational diversity. Theoretical analysis grounded in second-order neighborhood risk reveals the model’s invariance to generated perturbations and its curvature-suppression mechanism. Evaluated across seven benchmarks, the method consistently outperforms nine state-of-the-art augmentation techniques, achieving superior performance in classification, robustness, calibration, and transfer learning tasks across CNNs, Vision Transformers, and vision-language models.

0 citationsRead paper