Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用MéTRON-FR模型评估法语语法能力,通过特定任务测试发现模型在语法相关任务上表现良好,但在世界知识任务上表现不佳。
📝 Abstract
We submit MéTRON-FR, a 125M GPT-2 pretrained on 92.47M words of French, to the BabyLM 2026 Strict track. It scores 85.97 +/- 0.17% on QFrBLiMP (a native Quebec-French benchmark of grammatical minimal pairs) and 62.80% on the BabyLM-weighted leaderboard. A cross-lingual GLUE (General Language Understanding Evaluation) protocol that combines French task-data translation with rank-16 LoRA (Low-Rank Adaptation) produces a sharp task-type gradient: relational tasks gain measurably, while world-knowledge tasks regress. Bilingual Lexicon Induction aligns the French embeddings to GPT-2 at p@1 = 68.84 +/- 8.61%, 18X above chance, suggesting cross-lingual alignment tracks acquired grammatical competence rather than training duration. An ablation study shows that single-token zero-shot scoring is dominated by tokenizer and template artifacts at the child scale, motivating tokenizer-swap sensitivity, placebo-controlled prompting, and native-language minimal-pair benchmarks as standard diagnostics.
Problem

Research questions and friction points this paper is trying to address.

Native-Language Evaluation
Tokenizer Sensitivity
Cross-lingual Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

MéTRON-FR
tokenizer-swap sensitivity
cross-lingual GLUE
Bilingual Lexicon Induction
native-language minimal-pair benchmarks
🔎 Similar Papers
No similar papers found.