Institution profile

Delhi Technological University

Academic institutionasia · in
Official website
Research library35linked papers
Opportunities0open roles
Selected work

Representative Papers

LaMSUM: Amplifying Voices Against Harassment through LLM Guided Extractive Summarization of User Incident Reports

Jun 22, 2024

To address the challenge of manually reviewing large-scale, code-mixed sexual harassment reports in India’s Safe City platform, this paper proposes the first LLM-driven extractive summarization framework tailored to this domain. Methodologically, it introduces a multi-model collaborative architecture integrating Llama, Mistral, and GPT-4o, enhanced by hierarchical text segmentation, prompt-engineered fine-grained extraction decisions, and an ensemble voting mechanism—effectively mitigating LLMs’ abstraction bias and context window limitations. Contributions include: (1) the first explainable and traceable extractive summarization system for code-mixed harassment reports; (2) state-of-the-art performance on the Safe City dataset, significantly outperforming existing baselines; and (3) generation of high-fidelity, structured event overviews that directly inform evidence-based policymaking and targeted anti-harassment interventions.

1 citationsRead paper

HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection

Jul 17, 2026

This work addresses the key challenge in multimodal sarcasm and cyberbullying detection—modeling multi-level semantic incongruities between textual and visual modalities. To this end, the authors propose the HCIG framework, which introduces, for the first time, a hierarchical cross-modal inconsistency graph network. This architecture leverages graph attention mechanisms to capture fine-grained to holistic semantic contradictions across token, phrase, and global levels, and employs hierarchical attention to adaptively fuse multi-granularity representations. Additionally, a Graph-based Contradiction-aware Convolutional Network (GCCN) is designed with contradiction-aware graph pooling to enable efficient cross-modal reasoning. Experimental results demonstrate that HCIG achieves state-of-the-art performance, attaining 85.74% accuracy and 85.29% macro F1 on the MMSD dataset, and 69.62% accuracy with 74.90% bullying-class F1 on the MultiBully dataset.

0 citationsRead paper

LUMINA-26: Low-Light Understanding for Modeling and Interpreting Night-time Actions

Jun 22, 2026

This study addresses the challenges of human action recognition under low-light conditions, where insufficient illumination, noise interference, and motion blur severely degrade performance, compounded by the limited diversity and realism of existing datasets. To bridge this gap, the authors introduce LUMINA-26, the first large-scale real-world low-light action recognition dataset comprising 26 action classes and 6,784 videos. They further propose Illumi-Net, a novel illumination-adaptive mixture-of-experts network that integrates video-level illumination cues to jointly perform image enhancement and spatiotemporal feature extraction, enabling accurate recognition through conditional expert routing. The model achieves strong performance on ELLAR (Top-1: 55.13%, Top-5: 78.87%) and establishes a robust baseline on LUMINA-26 (Top-1: 75.95%, Top-5: 93.58%).

0 citationsRead paper
Recent publications

Latest Papers

HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection

Jul 17, 2026

This work addresses the key challenge in multimodal sarcasm and cyberbullying detection—modeling multi-level semantic incongruities between textual and visual modalities. To this end, the authors propose the HCIG framework, which introduces, for the first time, a hierarchical cross-modal inconsistency graph network. This architecture leverages graph attention mechanisms to capture fine-grained to holistic semantic contradictions across token, phrase, and global levels, and employs hierarchical attention to adaptively fuse multi-granularity representations. Additionally, a Graph-based Contradiction-aware Convolutional Network (GCCN) is designed with contradiction-aware graph pooling to enable efficient cross-modal reasoning. Experimental results demonstrate that HCIG achieves state-of-the-art performance, attaining 85.74% accuracy and 85.29% macro F1 on the MMSD dataset, and 69.62% accuracy with 74.90% bullying-class F1 on the MultiBully dataset.

0 citationsRead paper

LUMINA-26: Low-Light Understanding for Modeling and Interpreting Night-time Actions

Jun 22, 2026

This study addresses the challenges of human action recognition under low-light conditions, where insufficient illumination, noise interference, and motion blur severely degrade performance, compounded by the limited diversity and realism of existing datasets. To bridge this gap, the authors introduce LUMINA-26, the first large-scale real-world low-light action recognition dataset comprising 26 action classes and 6,784 videos. They further propose Illumi-Net, a novel illumination-adaptive mixture-of-experts network that integrates video-level illumination cues to jointly perform image enhancement and spatiotemporal feature extraction, enabling accurate recognition through conditional expert routing. The model achieves strong performance on ELLAR (Top-1: 55.13%, Top-5: 78.87%) and establishes a robust baseline on LUMINA-26 (Top-1: 75.95%, Top-5: 93.58%).

0 citationsRead paper

FlowFake: Liquid Networks for Audio Deepfake Detection

Jun 17, 2026

This work addresses the limited generalization of existing audio deepfake detection methods in cross-dataset scenarios, where performance degrades significantly on unseen synthetic speech. To tackle this challenge, the authors propose FlowFake, the first approach to integrate Liquid Time-constant (LTC) neural networks into deepfake detection. By modeling hidden states through learnable ordinary differential equations, FlowFake adaptively captures spectrotemporal and prosodic artifacts across multiple time scales. Remarkably, with only 34K parameters, FlowFake achieves up to 79.97% accuracy across four cross-domain benchmarks—matching the performance of models 300 times larger—and substantially enhances cross-dataset generalization under stringent parameter constraints.

0 citationsRead paper