The record is part of the task: matched-record evaluation of text classifiers across maintenance, safety and recall reporting
研究通过比较同一案例在不同阶段的记录,评估文本分类器在维护、安全和召回报告中的表现差异。
研究通过比较同一案例在不同阶段的记录,评估文本分类器在维护、安全和召回报告中的表现差异。
本文介绍了ImageEval 2026任务,通过多种方法如零样本提示、视觉-语言模型微调等解决阿拉伯语文化背景下的多模态评估问题。
为解决阿拉伯语自动语音识别面临的独特挑战,通过构建包含275名来自11个阿拉伯国家的说话者的多方言数据集BULBUL,并采用两层人工验证确保质量。
To address the scarcity of high-quality morphologically annotated corpora for the Arabic Qur’an, this study constructs the first manually curated, fine-grained morphological corpus—annotated by three linguistics experts—comprising 77,429 word tokens. Morphological analysis leverages the Qabas lexicon for precise lemmatization and adopts the SAMA/Qabas framework to annotate 40 fine-grained part-of-speech categories, all rigorously validated through expert adjudication. The key contribution lies in enabling structured interoperability with over 100 Arabic lexical resources and corpora—including Qabas—thereby substantially improving morphological parsing accuracy and resource reusability for religious texts. This open-source corpus has been integrated into the SinaLab platform, establishing a new authoritative benchmark for Arabic computational linguistics and Qur’anic text research.
研究通过比较同一案例在不同阶段的记录,评估文本分类器在维护、安全和召回报告中的表现差异。
本文介绍了ImageEval 2026任务,通过多种方法如零样本提示、视觉-语言模型微调等解决阿拉伯语文化背景下的多模态评估问题。
为解决阿拉伯语自动语音识别面临的独特挑战,通过构建包含275名来自11个阿拉伯国家的说话者的多方言数据集BULBUL,并采用两层人工验证确保质量。
To address the scarcity of high-quality morphologically annotated corpora for the Arabic Qur’an, this study constructs the first manually curated, fine-grained morphological corpus—annotated by three linguistics experts—comprising 77,429 word tokens. Morphological analysis leverages the Qabas lexicon for precise lemmatization and adopts the SAMA/Qabas framework to annotate 40 fine-grained part-of-speech categories, all rigorously validated through expert adjudication. The key contribution lies in enabling structured interoperability with over 100 Arabic lexical resources and corpora—including Qabas—thereby substantially improving morphological parsing accuracy and resource reusability for religious texts. This open-source corpus has been integrated into the SinaLab platform, establishing a new authoritative benchmark for Arabic computational linguistics and Qur’anic text research.