Institution profile

Paris-Sorbonne University Abu Dhabi

Academic institutionasia · ae
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

SPARK-IL: Spectral Retrieval-Augmented RAG for Knowledge-driven Deepfake Detection via Incremental Learning

Apr 04, 2026

This work addresses the limited generalization of existing AI-generated image detection methods to unseen generative models by proposing a novel incremental learning framework that integrates dual-path spectral analysis with retrieval-augmented generation (RAG). The approach employs four-band Fourier decomposition to extract frequency-domain features, combines a partially frozen ViT-L/14 encoder with Kolmogorov–Arnold Network (KAN)-based mixture-of-experts to model band-specific characteristics, and incorporates elastic weight consolidation for continual learning. Notably, it introduces, for the first time, a synergy between spectral consistency priors and RAG, leveraging a Milvus vector database for knowledge retrieval to enhance discriminative robustness. Evaluated on the UniversalFakeDetect benchmark encompassing 19 generative models, the method achieves an average accuracy of 94.6%, substantially outperforming current state-of-the-art techniques.

0 citationsRead paper

C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Object Detection

Aug 30, 2025

Fine-grained object detection—e.g., vehicle damage assessment—faces challenges from strong contextual dependencies and insufficient local feature modeling. To address this, we propose ContextDiff, a detection framework that jointly leverages global scene understanding and generative denoising. Methodologically, it adopts a conditional diffusion detection paradigm, incorporating a dedicated global context encoder and an end-to-end generative denoising training strategy. Its core innovation is a context-aware fusion module that employs cross-attention to dynamically integrate local proposal features with independently encoded global scene representations—thereby alleviating conventional conditional diffusion models’ overreliance on local features. Evaluated on the CarDD benchmark, ContextDiff achieves a 3.2% mAP improvement over prior state-of-the-art methods, establishing a new benchmark for fine-grained detection in complex, context-rich scenes.

0 citationsRead paper

CVPD at QIAS 2025 Shared Task: An Efficient Encoder-Based Approach for Islamic Inheritance Reasoning

Aug 30, 2025

To address the stringent accuracy requirements for heir identification and share calculation under Islamic inheritance law (ʿIlm al-Mawārīth), this paper proposes a lightweight, on-device discriminative framework. The core innovation is an Attentive Relevance Scoring mechanism that leverages Arabic-specific encoders—MARBERT, ArabicBERT, or AraBERT—to perform semantic relevance ranking over candidate heirs, thereby avoiding the computational overhead and output uncertainty inherent in generative models. Evaluated on the QIAS 2025 dataset, the MARBERT variant achieves 69.87% accuracy—slightly below API-based large language models like Gemini (87.6%) but with substantially reduced computational footprint, enabling local deployment and privacy-preserving applications. This work constitutes the first systematic application of discriminative semantic matching to Islamic inheritance reasoning, establishing a novel pathway for resource-constrained, religion-sensitive legal AI systems.

0 citationsRead paper

SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection

Jul 27, 2025

Natural scene text detection faces challenges including difficulty in recognizing multilingual scripts and arbitrarily shaped (e.g., curved) text, as well as insufficient semantic richness in visual cues. To address these, we propose a semantic-aware vision-language detection framework: it leverages the CLIP pre-trained model and an Asymptotic Feature Pyramid Network (AFPN) to construct multi-scale visual representations; introduces a text-to-pixel contrastive learning mechanism and a language-vision decoder that employs cross-attention for fine-grained cross-modal semantic alignment. This work is the first to incorporate strong linguistic semantics into end-to-end text detection, significantly enhancing robustness for complex scripts and curved text. Our method achieves state-of-the-art F-scores of 84.8% on MLT-2019 and 90.2% on CTW1500, surpassing prior approaches.

0 citationsRead paper
Recent publications

Latest Papers

SPARK-IL: Spectral Retrieval-Augmented RAG for Knowledge-driven Deepfake Detection via Incremental Learning

Apr 04, 2026

This work addresses the limited generalization of existing AI-generated image detection methods to unseen generative models by proposing a novel incremental learning framework that integrates dual-path spectral analysis with retrieval-augmented generation (RAG). The approach employs four-band Fourier decomposition to extract frequency-domain features, combines a partially frozen ViT-L/14 encoder with Kolmogorov–Arnold Network (KAN)-based mixture-of-experts to model band-specific characteristics, and incorporates elastic weight consolidation for continual learning. Notably, it introduces, for the first time, a synergy between spectral consistency priors and RAG, leveraging a Milvus vector database for knowledge retrieval to enhance discriminative robustness. Evaluated on the UniversalFakeDetect benchmark encompassing 19 generative models, the method achieves an average accuracy of 94.6%, substantially outperforming current state-of-the-art techniques.

0 citationsRead paper

C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Object Detection

Aug 30, 2025

Fine-grained object detection—e.g., vehicle damage assessment—faces challenges from strong contextual dependencies and insufficient local feature modeling. To address this, we propose ContextDiff, a detection framework that jointly leverages global scene understanding and generative denoising. Methodologically, it adopts a conditional diffusion detection paradigm, incorporating a dedicated global context encoder and an end-to-end generative denoising training strategy. Its core innovation is a context-aware fusion module that employs cross-attention to dynamically integrate local proposal features with independently encoded global scene representations—thereby alleviating conventional conditional diffusion models’ overreliance on local features. Evaluated on the CarDD benchmark, ContextDiff achieves a 3.2% mAP improvement over prior state-of-the-art methods, establishing a new benchmark for fine-grained detection in complex, context-rich scenes.

0 citationsRead paper

CVPD at QIAS 2025 Shared Task: An Efficient Encoder-Based Approach for Islamic Inheritance Reasoning

Aug 30, 2025

To address the stringent accuracy requirements for heir identification and share calculation under Islamic inheritance law (ʿIlm al-Mawārīth), this paper proposes a lightweight, on-device discriminative framework. The core innovation is an Attentive Relevance Scoring mechanism that leverages Arabic-specific encoders—MARBERT, ArabicBERT, or AraBERT—to perform semantic relevance ranking over candidate heirs, thereby avoiding the computational overhead and output uncertainty inherent in generative models. Evaluated on the QIAS 2025 dataset, the MARBERT variant achieves 69.87% accuracy—slightly below API-based large language models like Gemini (87.6%) but with substantially reduced computational footprint, enabling local deployment and privacy-preserving applications. This work constitutes the first systematic application of discriminative semantic matching to Islamic inheritance reasoning, establishing a novel pathway for resource-constrained, religion-sensitive legal AI systems.

0 citationsRead paper

SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection

Jul 27, 2025

Natural scene text detection faces challenges including difficulty in recognizing multilingual scripts and arbitrarily shaped (e.g., curved) text, as well as insufficient semantic richness in visual cues. To address these, we propose a semantic-aware vision-language detection framework: it leverages the CLIP pre-trained model and an Asymptotic Feature Pyramid Network (AFPN) to construct multi-scale visual representations; introduces a text-to-pixel contrastive learning mechanism and a language-vision decoder that employs cross-attention for fine-grained cross-modal semantic alignment. This work is the first to incorporate strong linguistic semantics into end-to-end text detection, significantly enhancing robustness for complex scripts and curved text. Our method achieves state-of-the-art F-scores of 84.8% on MLT-2019 and 90.2% on CTW1500, surpassing prior approaches.

0 citationsRead paper