Institution profile

Humain

Research institution
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding

Oct 19, 2025

Existing long-context evaluation benchmarks lack bilingual (English/Arabic) support and multitask design, making it difficult to rigorously assess LLMs’ deep reasoning, cross-document understanding, information tracking, and bilingual information extraction capabilities at context lengths of 4K–128K+ tokens. To address this gap, we propose BilingualLongEval—the first bilingual, multitask benchmark explicitly designed for long-context understanding. It comprises four challenging tasks: multi-document question answering, bilingual question answering, intra-paragraph claim verification, and long-text multiple-choice. The benchmark is built upon high-quality, manually curated and rigorously filtered bilingual data, emphasizing cross-lingual alignment, long-range dependency modeling, and logical consistency validation. Empirical evaluation reveals significant performance degradation across state-of-the-art models—including GPT-4o—demonstrating the benchmark’s high difficulty and effectiveness. BilingualLongEval thus establishes a new standard for evaluating both long-context reasoning and bilingual comprehension in LLMs.

0 citationsRead paper

Saudi Sign Language Translation Using T5

Oct 13, 2025

This work addresses the unique challenges of Saudi Sign Language (SSL) translation, particularly face occlusion prevalent in regional cultural contexts. To this end, we construct the first SSL parallel corpus featuring realistic occlusion scenarios and propose three hierarchical evaluation protocols to comprehensively assess robustness. Methodologically, we explore cross-lingual transfer learning based on the T5 architecture: pretraining on YouTubeASL followed by fine-tuning on SSL data. Experiments demonstrate that this strategy improves BLEU-4 by approximately threefold over a baseline trained solely on SSL data. Our key contributions are: (1) releasing the first SSL translation dataset annotated for facial occlusion and grounded in Middle Eastern sociocultural context; (2) providing the first empirical validation of effective cross-sign-language transfer from American Sign Language (ASL) to SSL, revealing cross-lingual representational transferability in sign language models; and (3) introducing a novel evaluation paradigm for occlusion-robust sign language translation.

0 citationsRead paper

A-SEA3L-QA: A Fully Automated Self-Evolving, Adversarial Workflow for Arabic Long-Context Question-Answer Generation

Sep 02, 2025

To address challenges in automatic Arabic long-text QA pair generation—including low output quality, uncontrollable difficulty levels, and the absence of standardized evaluation benchmarks—this paper proposes an end-to-end self-evolving adversarial multi-agent framework. Methodologically, it establishes a specialized large vision-language model (LVLM) collaboration system comprising a question generator, an answer ensemble, and an evaluator; integrates confidence-driven re-generation, closed-loop feedback, and automated data curation; and introduces a tunable, self-evolving difficulty mechanism. Key contributions include: (1) the first large-scale Arabic long-context evaluation benchmark, AraLongBench; (2) fully automated, human-free continual performance optimization; and (3) substantial improvements in long-context comprehension for mainstream Arabic LVLMs on AraLongBench, with superior QA pair quality and difficulty controllability compared to static pipeline approaches.

0 citationsRead paper

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality

Jul 27, 2025

To address performance degradation in vision-language models (VLMs) caused by image-text misalignment and noisy training data, this paper proposes a lightweight, self-consistent data filtering framework. The method leverages only a fine-tuned small-scale VLM—without auxiliary modules or large language/model guidance—to jointly assess image quality, text fluency, and cross-modal semantic alignment. It employs a context-aware, end-to-end judgment mechanism, enabling efficient data purification with minimal computational overhead. Experiments demonstrate that datasets filtered by our framework substantially outperform the original noisy datasets across multiple downstream vision-language tasks—and even rival human-annotated high-quality benchmarks. These results validate the effectiveness and practicality of the novel paradigm “leveraging small VLMs to drive high-quality dataset construction.”

0 citationsRead paper

Multi-Agent Interactive Question Generation Framework for Long Document Understanding

Jul 27, 2025

Document understanding (DU) in long-context and complex-layout scenarios remains hindered by the scarcity of fine-grained annotations—particularly for low-resource languages like Arabic, which heavily rely on costly manual labeling. To address this, we propose the first fully automated multi-agent interaction framework that synergistically integrates structured prompting, cross-lingual question generation, and layout-aware mechanisms to efficiently synthesize high-quality single-page and multi-page QA pairs in English and Arabic. The resulting benchmark, AraEngLongBench, exhibits rich linguistic diversity, intricate document layouts, and substantial long-context challenges, imposing rigorous evaluation pressure on both mainstream open-source and proprietary vision-language models. Our approach breaks from traditional manual annotation paradigms, significantly improving scalability, fidelity, and efficiency in long-document DU data synthesis. It establishes a viable, extensible technical pathway for advancing DU in low-resource languages.

0 citationsRead paper
Recent publications

Latest Papers

LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding

Oct 19, 2025

Existing long-context evaluation benchmarks lack bilingual (English/Arabic) support and multitask design, making it difficult to rigorously assess LLMs’ deep reasoning, cross-document understanding, information tracking, and bilingual information extraction capabilities at context lengths of 4K–128K+ tokens. To address this gap, we propose BilingualLongEval—the first bilingual, multitask benchmark explicitly designed for long-context understanding. It comprises four challenging tasks: multi-document question answering, bilingual question answering, intra-paragraph claim verification, and long-text multiple-choice. The benchmark is built upon high-quality, manually curated and rigorously filtered bilingual data, emphasizing cross-lingual alignment, long-range dependency modeling, and logical consistency validation. Empirical evaluation reveals significant performance degradation across state-of-the-art models—including GPT-4o—demonstrating the benchmark’s high difficulty and effectiveness. BilingualLongEval thus establishes a new standard for evaluating both long-context reasoning and bilingual comprehension in LLMs.

0 citationsRead paper

Saudi Sign Language Translation Using T5

Oct 13, 2025

This work addresses the unique challenges of Saudi Sign Language (SSL) translation, particularly face occlusion prevalent in regional cultural contexts. To this end, we construct the first SSL parallel corpus featuring realistic occlusion scenarios and propose three hierarchical evaluation protocols to comprehensively assess robustness. Methodologically, we explore cross-lingual transfer learning based on the T5 architecture: pretraining on YouTubeASL followed by fine-tuning on SSL data. Experiments demonstrate that this strategy improves BLEU-4 by approximately threefold over a baseline trained solely on SSL data. Our key contributions are: (1) releasing the first SSL translation dataset annotated for facial occlusion and grounded in Middle Eastern sociocultural context; (2) providing the first empirical validation of effective cross-sign-language transfer from American Sign Language (ASL) to SSL, revealing cross-lingual representational transferability in sign language models; and (3) introducing a novel evaluation paradigm for occlusion-robust sign language translation.

0 citationsRead paper

A-SEA3L-QA: A Fully Automated Self-Evolving, Adversarial Workflow for Arabic Long-Context Question-Answer Generation

Sep 02, 2025

To address challenges in automatic Arabic long-text QA pair generation—including low output quality, uncontrollable difficulty levels, and the absence of standardized evaluation benchmarks—this paper proposes an end-to-end self-evolving adversarial multi-agent framework. Methodologically, it establishes a specialized large vision-language model (LVLM) collaboration system comprising a question generator, an answer ensemble, and an evaluator; integrates confidence-driven re-generation, closed-loop feedback, and automated data curation; and introduces a tunable, self-evolving difficulty mechanism. Key contributions include: (1) the first large-scale Arabic long-context evaluation benchmark, AraLongBench; (2) fully automated, human-free continual performance optimization; and (3) substantial improvements in long-context comprehension for mainstream Arabic LVLMs on AraLongBench, with superior QA pair quality and difficulty controllability compared to static pipeline approaches.

0 citationsRead paper

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality

Jul 27, 2025

To address performance degradation in vision-language models (VLMs) caused by image-text misalignment and noisy training data, this paper proposes a lightweight, self-consistent data filtering framework. The method leverages only a fine-tuned small-scale VLM—without auxiliary modules or large language/model guidance—to jointly assess image quality, text fluency, and cross-modal semantic alignment. It employs a context-aware, end-to-end judgment mechanism, enabling efficient data purification with minimal computational overhead. Experiments demonstrate that datasets filtered by our framework substantially outperform the original noisy datasets across multiple downstream vision-language tasks—and even rival human-annotated high-quality benchmarks. These results validate the effectiveness and practicality of the novel paradigm “leveraging small VLMs to drive high-quality dataset construction.”

0 citationsRead paper

Multi-Agent Interactive Question Generation Framework for Long Document Understanding

Jul 27, 2025

Document understanding (DU) in long-context and complex-layout scenarios remains hindered by the scarcity of fine-grained annotations—particularly for low-resource languages like Arabic, which heavily rely on costly manual labeling. To address this, we propose the first fully automated multi-agent interaction framework that synergistically integrates structured prompting, cross-lingual question generation, and layout-aware mechanisms to efficiently synthesize high-quality single-page and multi-page QA pairs in English and Arabic. The resulting benchmark, AraEngLongBench, exhibits rich linguistic diversity, intricate document layouts, and substantial long-context challenges, imposing rigorous evaluation pressure on both mainstream open-source and proprietary vision-language models. Our approach breaks from traditional manual annotation paradigms, significantly improving scalability, fidelity, and efficiency in long-document DU data synthesis. It establishes a viable, extensible technical pathway for advancing DU in low-resource languages.

0 citationsRead paper