🤖 AI Summary
Existing evaluations of large language models (LLMs) lack rigorous assessment of grammatical comprehension in Quebec French, a low-resource regional variety. Method: We introduce the first fine-grained, human-curated benchmark for Quebec French grammar—comprising 1,761 minimal-pair items spanning 20 grammatical phenomena, derived from Canadian government corpora and validated by 12 native speakers. We systematically evaluate LLMs of varying scales using probabilistic scoring. Results: While model performance improves with parameter count, all models exhibit substantial deficits on grammatical tasks requiring deep semantic-syntactic integration, significantly underperforming human annotators. This work establishes the first interpretable, reproducible evaluation framework for regional French varieties, revealing systematic weaknesses at the semantics–syntax interface in current LLMs. It further provides a methodological blueprint for assessing grammatical competence in other under-resourced dialects.
📝 Abstract
In this paper, we introduce the Quebec-French Benchmark of Linguistic Minimal Pairs (QFrBLiMP), a corpus designed to evaluate the linguistic knowledge of LLMs on prominent grammatical phenomena in Quebec-French. QFrBLiMP consists of 1,761 minimal pairs annotated with 20 linguistic phenomena. Specifically, these minimal pairs have been created by manually modifying sentences extracted from an official online resource maintained by a Québec government institution. Each pair is annotated by twelve Quebec-French native speakers, who select the sentence they feel is grammatical amongst the two. These annotations are used to compare the competency of LLMs with that of humans. We evaluate different LLMs on QFrBLiMP and MultiBLiMP-Fr by observing the rate of higher probabilities assigned to the sentences of each minimal pair for each category. We find that while grammatical competence scales with model size, a clear hierarchy of difficulty emerges. All benchmarked models consistently fail on phenomena requiring deep semantic understanding, revealing a critical limitation and a significant gap compared to human performance on these specific tasks.