đ€ AI Summary
Existing NLP evaluation benchmarks lack standardized, high-quality resources for assessing large language modelsâ (LLMs) comprehension of regionally specific linguistic expressionsâparticularly in low-resource dialects such as Quebec French. Method: We introduce two rigorously constructed, reproducible benchmark datasetsâQFrCoRE (4,633 regionally attested idioms) and QFrCoRT (171 regionally marked lexical items)âand propose a novel âidiom-as-probeâ paradigm for dialect capability assessment. Leveraging systematic corpus collection and expert annotation, we conduct empirical evaluation across 94 LLMs. Contribution/Results: Our benchmarks effectively discriminate fine-grained semantic understanding of Quebec French, revealing substantial performance gaps across models. They provide the first reliable, granular evaluation framework for dialectal adaptation, addressing a critical gap in low-resource dialect comprehension assessment and enabling targeted development of geographically aware language models.
đ Abstract
The tasks of idiom understanding and dialect understanding are both well-established benchmarks in natural language processing. In this paper, we propose combining them, and using regional idioms as a test of dialect understanding. Towards this end, we propose two new benchmark datasets for the Quebec dialect of French: QFrCoRE, which contains 4,633 instances of idiomatic phrases, and QFrCoRT, which comprises 171 regional instances of idiomatic words. We explain how to construct these corpora, so that our methodology can be replicated for other dialects. Our experiments with 94 LLM demonstrate that our regional idiom benchmarks are a reliable tool for measuring a model's proficiency in a specific dialect.