Institution profile

M42 Health

Industry researchasia · ae
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation

Jan 27, 2026

Existing text generation evaluation metrics, such as BLEU and BERTScore, struggle to effectively assess semantic fidelity and often overlook critical errors like content omissions or factual inconsistencies. This work proposes a reference-free, multidimensional evaluation framework that introduces a cross-examination mechanism into generation assessment for the first time: treating the source and generated texts as independent knowledge bases, it performs question-answer–based mutual validation by generating verifiable questions from each to interrogate the other. The method yields three interpretable scores—coverage, consistency, and conformity—without requiring reference texts, enabling precise detection of semantic distortions. It significantly outperforms conventional metrics across translation, summarization, and clinical note tasks, effectively capturing errors at the entity and relational levels. Its reference-free and reference-based modes exhibit strong correlation, and expert validation confirms that mismatched questions align closely with actual semantic errors.

0 citationsRead paper

Bridging Language Barriers in Healthcare: A Study on Arabic LLMs

Jan 16, 2025

Translation-based data augmentation fails to guarantee clinical performance for target languages in multilingual medical AI, particularly for low-resource languages like Arabic. Method: We propose a task-aware language-mixing sampling strategy for Arabic large language models, identifying task-specific optimal language proportions; conduct multilingual pretraining coupled with ablation-driven joint analysis of data composition and scale to assess scalability and efficacy. Contribution/Results: Empirical evaluation on clinical benchmarks—including diagnostic reasoning and clinical note generation—demonstrates that optimized language mixing improves Arabic clinical accuracy by 12.4% and significantly enhances model robustness. Our work establishes a reproducible, methodology-driven framework for data composition optimization in low-resource medical language modeling, offering both principled guidelines and practical strategies for multilingual clinical LLM development.

0 citationsRead paper
Recent publications

Latest Papers

Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation

Jan 27, 2026

Existing text generation evaluation metrics, such as BLEU and BERTScore, struggle to effectively assess semantic fidelity and often overlook critical errors like content omissions or factual inconsistencies. This work proposes a reference-free, multidimensional evaluation framework that introduces a cross-examination mechanism into generation assessment for the first time: treating the source and generated texts as independent knowledge bases, it performs question-answer–based mutual validation by generating verifiable questions from each to interrogate the other. The method yields three interpretable scores—coverage, consistency, and conformity—without requiring reference texts, enabling precise detection of semantic distortions. It significantly outperforms conventional metrics across translation, summarization, and clinical note tasks, effectively capturing errors at the entity and relational levels. Its reference-free and reference-based modes exhibit strong correlation, and expert validation confirms that mismatched questions align closely with actual semantic errors.

0 citationsRead paper

Bridging Language Barriers in Healthcare: A Study on Arabic LLMs

Jan 16, 2025

Translation-based data augmentation fails to guarantee clinical performance for target languages in multilingual medical AI, particularly for low-resource languages like Arabic. Method: We propose a task-aware language-mixing sampling strategy for Arabic large language models, identifying task-specific optimal language proportions; conduct multilingual pretraining coupled with ablation-driven joint analysis of data composition and scale to assess scalability and efficacy. Contribution/Results: Empirical evaluation on clinical benchmarks—including diagnostic reasoning and clinical note generation—demonstrates that optimized language mixing improves Arabic clinical accuracy by 12.4% and significantly enhances model robustness. Our work establishes a reproducible, methodology-driven framework for data composition optimization in low-resource medical language modeling, offering both principled guidelines and practical strategies for multilingual clinical LLM development.

0 citationsRead paper