MGAL: A Multilingual Granularity-Aware Long-Context Benchmark

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决长上下文大语言模型在细粒度和多语言评估上的不足,本文提出了MGAL基准,通过联合国报告构建了跨六种语言的多层次、位置敏感数据集,并进行了系统性分析。
📝 Abstract
Evaluation of long-context Large Language Models (LLMs) has advanced rapidly. However, most existing benchmarks are limited to the document level and focus mainly on high-resource languages, leaving many fine-grained challenges insufficiently evaluated. To address this gap, we present MGAL, the first multilingual, granularity- and position-aware long-context benchmark. MGAL is constructed from United Nations (UN) reports spanning 8K to 128K tokens across the six official UN languages. It covers four coherent levels of linguistic granularity (word, sentence, paragraph, and document) and further stratifies entries by their position within the document (begin, middle, and end), indexed at both the document and paragraph levels. This design enables systematic diagnosis of multilingual long-context comprehension across different granularities. Through extensive experiments and analyses, we find that: (1) LLMs perform well at word-level tasks but struggle with coarser-grained ones; and (2) Closed-source models retain a clear performance advantage in lower-resource languages. We further identify two new challenges: (1) Under local semantic crowding, where neighboring sentences share topics and entities, models tend to follow surface cues (e.g., connectives like ``however'' or repeated entities) rather than the discourse role of the sentence in surrounding context (e.g., background, outcome); and (2) A gap between fluency and consistency in generated outputs, where models produce text that reads smoothly but drifts from the source facts. In addition, we observe several patterns in line with prior studies, including reliance on nearby evidence and reuse of options under uncertainty.
Problem

Research questions and friction points this paper is trying to address.

long-context
multilingual
granularity-aware
position-aware
semantic crowding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multilingual
Granularity-Aware
Long-Context Benchmark
Position-Aware
Semantic Crowding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chunhan Li
The Hong Kong University of Science and Technology (Guangzhou)
Chenglin Xu
Chenglin Xu
China University of Petroleum (East China)
Z
Zongyang Zhang
China University of Petroleum (East China)
J
Jiale Liu
China University of Petroleum (East China)
Z
Zhuoxi Rao
Northeastern University at Qinhuangdao
X
Xudong Jia
China University of Petroleum (East China)
J
Junxiu He
China University of Petroleum (East China)
Menglin Yang
Menglin Yang
HKUST(GZ) | Yale University | CUHK
Hyperbolic Representation LearningTransformerRecommender SystemLLM
W
Wenjuan Gong
China University of Petroleum (East China)
Zhengzhe Liu
Zhengzhe Liu
Lingnan University
Computer Vision3D GenerationComputer GraphicsAI AgentAI4Sci
Chengwei Qin
Chengwei Qin
HKUST(GZ), NTU
LLMNLP