Institution profile

University of Jordan

Academic institutionasia · jo
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

TellTale: Blending Multi-Instance LoRA Text Encoders and a Zero-Shot LLM Judge for Ambivalence/Hesitancy Recognition in Videos

Jul 18, 2026

This work presents the first approach to achieve high-performance recognition of speaker contradiction and hesitation states using only interview transcript text, without relying on audio or visual modalities. The method integrates two LoRA-enhanced multilingual text encoders—multilingual-e5-large and mDeBERTa-v3-base—fine-tuned via multiple instance learning, along with a quantized 14B instruction-tuned large language model employed as a zero-shot scorer. Video-level predictions are derived through weighted averaging and a unified threshold determined by group cross-validation. Evaluated on the private test set of the BAH dataset, the approach attains a Macro-F1 score of 0.7364 and an average precision of 0.7940, substantially outperforming the official vision-based baseline (Macro-F1: 0.2827), thereby demonstrating the effectiveness and novelty of a purely text-driven solution.

0 citationsRead paper

Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation

Jun 03, 2025Computer Science Review

To address critical challenges in multilingual image captioning—including semantic inconsistency across non-English languages, data scarcity, and weak cross-lingual reasoning—this paper systematically evaluates the generalization capability of attention-based Transformer models. Methodologically, it integrates multi-head self-attention, cross-modal alignment, and zero-shot/few-shot transfer learning, jointly leveraging multilingual BERT and XLM-R for language-agnostic visual–textual encoding. The work introduces two key contributions: (1) the first unified cross-lingual image captioning benchmark, and (2) a language-agnostic attention interpretability framework for probing cross-lingual alignment. Extensive experiments across 12 languages and 5 datasets demonstrate an average BLEU-4 improvement of 9.2%, with non-English captions achieving 87% fluency relative to native speakers. The approach significantly enhances cross-lingual consistency and model robustness.

0 citationsRead paper

Parameter Efficient Fine Tuning Llama 3.1 for Answering Arabic Legal Questions: A Case Study on Jordanian Laws

Apr 28, 20252025 1st International Conference on Computational Intelligence Approaches and Applications (ICCIAA)

This study addresses the poor adaptability and high computational cost of large language models in Arabic legal question answering, particularly for jurisdiction-specific contexts such as Jordanian law. To this end, the authors construct the first structured question-answering dataset focused on Jordanian law, comprising 6,000 samples. They further propose an efficient fine-tuning approach that uniquely combines parameter-efficient fine-tuning (PEFT) with 4-bit quantization, leveraging the Llama-3.1-8B model enhanced with LoRA adapters and optimized via the Unsloth framework. Experimental results demonstrate that this method substantially reduces computational resource requirements while significantly improving the model’s reasoning capabilities and answer accuracy in Arabic legal QA tasks. The effectiveness of the approach is corroborated by BLEU and ROUGE evaluation metrics.

0 citationsRead paper

An ensemble model with attention based mechanism for image captioning

Jan 22, 2025Computers & electrical engineering

This work addresses image captioning, aiming to enhance visual understanding and improve descriptive accuracy—particularly for long-tail visual scenes. We propose a multi-encoder–decoder architecture that jointly integrates CNN, LSTM, and Transformer encoders, coupled with a hierarchical attention mechanism to dynamically align visual regions with semantic tokens. To our knowledge, we are the first to introduce a learnable cross-modal ensemble weighting strategy, embedding adaptive attention gating into a weighted ensemble decoding framework. Furthermore, reinforcement learning is employed for post-hoc fine-tuning to optimize caption quality. Evaluated on the COCO benchmark, our method achieves BLEU-4 = 38.7 and CIDEr = 128.4—substantially outperforming both individual models and conventional ensemble approaches. These results validate the effectiveness of our fine-grained visual–linguistic alignment and robust, adaptive multimodal integration.

0 citationsRead paper
Recent publications

Latest Papers

TellTale: Blending Multi-Instance LoRA Text Encoders and a Zero-Shot LLM Judge for Ambivalence/Hesitancy Recognition in Videos

Jul 18, 2026

This work presents the first approach to achieve high-performance recognition of speaker contradiction and hesitation states using only interview transcript text, without relying on audio or visual modalities. The method integrates two LoRA-enhanced multilingual text encoders—multilingual-e5-large and mDeBERTa-v3-base—fine-tuned via multiple instance learning, along with a quantized 14B instruction-tuned large language model employed as a zero-shot scorer. Video-level predictions are derived through weighted averaging and a unified threshold determined by group cross-validation. Evaluated on the private test set of the BAH dataset, the approach attains a Macro-F1 score of 0.7364 and an average precision of 0.7940, substantially outperforming the official vision-based baseline (Macro-F1: 0.2827), thereby demonstrating the effectiveness and novelty of a purely text-driven solution.

0 citationsRead paper

Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation

Jun 03, 2025Computer Science Review

To address critical challenges in multilingual image captioning—including semantic inconsistency across non-English languages, data scarcity, and weak cross-lingual reasoning—this paper systematically evaluates the generalization capability of attention-based Transformer models. Methodologically, it integrates multi-head self-attention, cross-modal alignment, and zero-shot/few-shot transfer learning, jointly leveraging multilingual BERT and XLM-R for language-agnostic visual–textual encoding. The work introduces two key contributions: (1) the first unified cross-lingual image captioning benchmark, and (2) a language-agnostic attention interpretability framework for probing cross-lingual alignment. Extensive experiments across 12 languages and 5 datasets demonstrate an average BLEU-4 improvement of 9.2%, with non-English captions achieving 87% fluency relative to native speakers. The approach significantly enhances cross-lingual consistency and model robustness.

0 citationsRead paper

Parameter Efficient Fine Tuning Llama 3.1 for Answering Arabic Legal Questions: A Case Study on Jordanian Laws

Apr 28, 20252025 1st International Conference on Computational Intelligence Approaches and Applications (ICCIAA)

This study addresses the poor adaptability and high computational cost of large language models in Arabic legal question answering, particularly for jurisdiction-specific contexts such as Jordanian law. To this end, the authors construct the first structured question-answering dataset focused on Jordanian law, comprising 6,000 samples. They further propose an efficient fine-tuning approach that uniquely combines parameter-efficient fine-tuning (PEFT) with 4-bit quantization, leveraging the Llama-3.1-8B model enhanced with LoRA adapters and optimized via the Unsloth framework. Experimental results demonstrate that this method substantially reduces computational resource requirements while significantly improving the model’s reasoning capabilities and answer accuracy in Arabic legal QA tasks. The effectiveness of the approach is corroborated by BLEU and ROUGE evaluation metrics.

0 citationsRead paper

An ensemble model with attention based mechanism for image captioning

Jan 22, 2025Computers & electrical engineering

This work addresses image captioning, aiming to enhance visual understanding and improve descriptive accuracy—particularly for long-tail visual scenes. We propose a multi-encoder–decoder architecture that jointly integrates CNN, LSTM, and Transformer encoders, coupled with a hierarchical attention mechanism to dynamically align visual regions with semantic tokens. To our knowledge, we are the first to introduce a learnable cross-modal ensemble weighting strategy, embedding adaptive attention gating into a weighted ensemble decoding framework. Furthermore, reinforcement learning is employed for post-hoc fine-tuning to optimize caption quality. Evaluated on the COCO benchmark, our method achieves BLEU-4 = 38.7 and CIDEr = 128.4—substantially outperforming both individual models and conventional ensemble approaches. These results validate the effectiveness of our fine-grained visual–linguistic alignment and robust, adaptive multimodal integration.

0 citationsRead paper