Institution profile

Qatar Computing Research Institute

Academic institutionasia · qa
Official website
Research library137linked papers
Opportunities0open roles
Selected work

Representative Papers

AIMA at SemEval-2024 Task 10: History-Based Emotion Recognition in Hindi-English Code-Mixed Conversations

Jan 19, 2025International Workshop on Semantic Evaluation

This work addresses emotion recognition in Hindi-English code-mixed (Hinglish) dialogues. We propose a history-aware multimodal framework comprising two key components: (1) a Hinglish-to-English translation pre-processing pipeline for linguistic normalization, and (2) a joint contextual modeling architecture integrating bidirectional LSTM or Transformer-based context encoders with an ensemble of multilingual pretrained language models (BERT, RoBERTa, XLM-R). Crucially, we introduce the novel concept of “history-aware contextual modeling”, synergistically coupled with code-mixed translation pre-processing. This design enhances cross-lingual robustness—particularly critical in low-resource emotion recognition in conversations (ERC). Evaluated on SemEval-2024 Task 10 Subtask 1, our approach outperforms all baseline systems, demonstrating the efficacy of jointly leveraging contextual awareness and language normalization for code-mixed emotion classification.

2 citationsRead paper

SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?

Feb 03, 2026

This work addresses the limited spatial reasoning capabilities of current vision-language models (VLMs) in real-world, unconstrained scenarios, where visual noise and diverse spatial relationships pose significant challenges. To this end, we introduce SpatiaLab—the first comprehensive benchmark for spatial reasoning in natural settings—encompassing 30 fine-grained tasks across six categories: relative position, depth occlusion, orientation, scale, navigation, and 3D geometry. The benchmark includes 1,400 real-world visual question-answering pairs and supports both multiple-choice and open-ended evaluation formats. Systematic evaluations of leading open- and closed-source, general-purpose and specialized VLMs reveal a substantial performance gap compared to human capabilities—for instance, InternVL3.5-72B achieves only 54.93% accuracy on multiple-choice questions versus 87.57% for humans—highlighting critical bottlenecks in complex spatial understanding.

1 citationsRead paper

Harmonizing the Arabic Audio Space with Data Scheduling

Jan 18, 2026

This study addresses the limited adaptability of large audio language models to Arabic’s diverse dialects and complex linguistic structures, as well as the lack of unified optimization across generative and discriminative tasks. To this end, we present the first systematic investigation into multitask training of large audio models for Arabic, introducing AraMega-SSum—the first Arabic speech summarization dataset—and proposing two novel strategies based on Qwen2.5-Omni (7B): Task-Progressive Curriculum (TPC) and Aligner-based Diverse Sampling (ADS). Our approach significantly accelerates early-stage convergence, improves F1 scores for paralinguistic attribute recognition, and stabilizes decoding in generative tasks. This work offers an efficient solution for adapting Omni-style models to low-resource, multimodal scenarios involving linguistically complex and underrepresented languages such as Arabic.

1 citationsRead paper

HCT-QA: A Benchmark for Question Answering on Human-Centric Tables

Mar 09, 2025arXiv.org

Human-centered tables (HCTs) exhibit complex layouts and heterogeneous formats, rendering existing data extraction and querying methods inadequate for effective question answering. To address this, we introduce HCT-QA—the first dedicated QA benchmark for HCTs—comprising over 6,800 real-world and synthetically generated tables and 77K natural-language question-answer pairs. We formally define the HCT QA task and propose a hybrid construction methodology integrating real-document parsing (from PDF/HTML), human verification, and controllable synthetic generation—moving beyond SQL-centric paradigms to enable LLM-native table understanding. We conduct zero-shot and few-shot evaluations across leading open- and closed-source LLMs, revealing an average F1 score below 40%, exposing critical limitations in layout awareness and cross-cell reasoning. HCT-QA establishes a new standard for evaluating and advancing HCT comprehension models.

1 citationsRead paper

A Perspective on Symbolic Machine Learning in Physical Sciences

Feb 25, 2025

The opacity of deep neural networks severely limits the applicability of machine learning in physical sciences, where interpretability, derivability, and pattern recognition are equally essential. Method: We propose Symbolic Machine Learning (Symbolic ML) as a complementary paradigm to numerical ML, systematically integrating symbolic regression, program induction, logical reasoning, and domain-knowledge embedding to transcend the black-box paradigm. Contribution/Results: We establish, for the first time, the epistemologically equal and synergistic roles of symbolic and numerical ML in physics; clarify fundamental methodological distinctions between ML and traditional physical reasoning; and formulate a new physics-aware AI paradigm—characterized by interpretability, verifiability, and generalizability. This framework provides both theoretical foundations and practical guidelines for accelerating scientific discovery through intelligible, principled, and reproducible AI.

1 citationsRead paper
Recent publications

Latest Papers