LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding

📅 2025-10-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing long-context evaluation benchmarks lack bilingual (English/Arabic) support and multitask design, making it difficult to rigorously assess LLMs’ deep reasoning, cross-document understanding, information tracking, and bilingual information extraction capabilities at context lengths of 4K–128K+ tokens. To address this gap, we propose BilingualLongEval—the first bilingual, multitask benchmark explicitly designed for long-context understanding. It comprises four challenging tasks: multi-document question answering, bilingual question answering, intra-paragraph claim verification, and long-text multiple-choice. The benchmark is built upon high-quality, manually curated and rigorously filtered bilingual data, emphasizing cross-lingual alignment, long-range dependency modeling, and logical consistency validation. Empirical evaluation reveals significant performance degradation across state-of-the-art models—including GPT-4o—demonstrating the benchmark’s high difficulty and effectiveness. BilingualLongEval thus establishes a new standard for evaluating both long-context reasoning and bilingual comprehension in LLMs.

Technology Category

Application Category

📝 Abstract
Recent advancements in Large Language Models (LLMs) have demonstrated sophisticated capabilities, including the ability to process and comprehend extended contexts. These emergent capabilities necessitate rigorous evaluation methods to effectively assess their performance in long-context understanding. In this paper, we present extbf{LC-Eval}, a bilingual, multi-task evaluation benchmark designed to evaluate long-context understanding in English and Arabic, targeting context lengths ranging from 4k to over 128k tokens. LC-Eval introduces four novel and challenging tasks: multi-document question answering, bilingual question answering, claim verification within a paragraph, and multiple-choice questions based on long contexts. These tasks are designed to assess LLMs' abilities in deep reasoning, document comprehension, information tracing, and bilingual information extraction and understanding. The benchmark includes datasets in both Arabic and English for each task, allowing for a comparative analysis of their performance across different text genres. Evaluations were conducted on both open-weight and closed LLMs, with results indicating that LC-Eval presents significant challenges. Even high-performing models, such as GPT-4o, struggled with certain tasks, highlighting the complexity and rigor of the benchmark.
Problem

Research questions and friction points this paper is trying to address.

Evaluating long-context understanding in English and Arabic
Assessing LLMs' reasoning and comprehension over extended texts
Benchmarking bilingual information extraction from lengthy documents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Developed bilingual multi-task benchmark for evaluation
Introduced four novel tasks for long-context assessment
Created datasets in both English and Arabic languages
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Sheikh Jubair
Sheikh Jubair
HUMAIN
A
Arwa Omayrah
HUMAIN
A
Amal Alshammari
Saudi Data and AI Authority
Alhanoof Althnian
Alhanoof Althnian
HUMAIN
A
Abdulhamed Alothaimen
HUMAIN
N
Norah A. Alzahrani
HUMAIN
S
Shahad D. Alzaidi
Saudi Data and AI Authority
N
Nora Al-Twairesh
HUMAIN, King Saud University
Abdulmohsen Al-Thubaity
Abdulmohsen Al-Thubaity
Principal AI Researcher, HUMAIN
Large Language Models (Data and Evaluation)Arabic Natural Language ProcessingComputational Lingu