LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding
Existing long-context evaluation benchmarks lack bilingual (English/Arabic) support and multitask design, making it difficult to rigorously assess LLMs’ deep reasoning, cross-document understanding, information tracking, and bilingual information extraction capabilities at context lengths of 4K–128K+ tokens. To address this gap, we propose BilingualLongEval—the first bilingual, multitask benchmark explicitly designed for long-context understanding. It comprises four challenging tasks: multi-document question answering, bilingual question answering, intra-paragraph claim verification, and long-text multiple-choice. The benchmark is built upon high-quality, manually curated and rigorously filtered bilingual data, emphasizing cross-lingual alignment, long-range dependency modeling, and logical consistency validation. Empirical evaluation reveals significant performance degradation across state-of-the-art models—including GPT-4o—demonstrating the benchmark’s high difficulty and effectiveness. BilingualLongEval thus establishes a new standard for evaluating both long-context reasoning and bilingual comprehension in LLMs.