🤖 AI Summary
Existing long-context evaluation benchmarks lack bilingual (English/Arabic) support and multitask design, making it difficult to rigorously assess LLMs’ deep reasoning, cross-document understanding, information tracking, and bilingual information extraction capabilities at context lengths of 4K–128K+ tokens. To address this gap, we propose BilingualLongEval—the first bilingual, multitask benchmark explicitly designed for long-context understanding. It comprises four challenging tasks: multi-document question answering, bilingual question answering, intra-paragraph claim verification, and long-text multiple-choice. The benchmark is built upon high-quality, manually curated and rigorously filtered bilingual data, emphasizing cross-lingual alignment, long-range dependency modeling, and logical consistency validation. Empirical evaluation reveals significant performance degradation across state-of-the-art models—including GPT-4o—demonstrating the benchmark’s high difficulty and effectiveness. BilingualLongEval thus establishes a new standard for evaluating both long-context reasoning and bilingual comprehension in LLMs.
📝 Abstract
Recent advancements in Large Language Models (LLMs) have demonstrated sophisticated capabilities, including the ability to process and comprehend extended contexts. These emergent capabilities necessitate rigorous evaluation methods to effectively assess their performance in long-context understanding. In this paper, we present extbf{LC-Eval}, a bilingual, multi-task evaluation benchmark designed to evaluate long-context understanding in English and Arabic, targeting context lengths ranging from 4k to over 128k tokens. LC-Eval introduces four novel and challenging tasks: multi-document question answering, bilingual question answering, claim verification within a paragraph, and multiple-choice questions based on long contexts. These tasks are designed to assess LLMs' abilities in deep reasoning, document comprehension, information tracing, and bilingual information extraction and understanding. The benchmark includes datasets in both Arabic and English for each task, allowing for a comparative analysis of their performance across different text genres. Evaluations were conducted on both open-weight and closed LLMs, with results indicating that LC-Eval presents significant challenges. Even high-performing models, such as GPT-4o, struggled with certain tasks, highlighting the complexity and rigor of the benchmark.