Which Forms of Caregiver Feedback Support Grammar Learning? A Reinforcement-Learning Study of Child-Like Language Models
研究使用儿童式语言模型通过强化学习测试不同形式的反馈对语法学习的支持作用,发现结构对齐反馈最能提升语法正确性。
研究使用儿童式语言模型通过强化学习测试不同形式的反馈对语法学习的支持作用,发现结构对齐反馈最能提升语法正确性。
研究解决了无监督语音模型对不同口音适应的问题,通过引入ABX-Accent基准和使用自适应域归一化方法微调预训练模型来改善跨说话者ABX得分。
研究使用检索增强语言模型测试情景记忆机制是否能缩小词汇频率对句法对比敏感度的影响,发现该方法有效但未完全解决问题。
This study investigates whether political polarization observed on social media accurately reflects the broader public’s attitudes, focusing on the ideological structure in France within a context of issue misalignment. By integrating large-scale X/Twitter data with nationally representative survey responses and employing dimensionality reduction alongside hierarchical modeling, the research examines how structural factors—such as user activity and visibility—influence the expression of political attitudes. The findings reveal that both online and offline political orientations consistently align along two stable dimensions: “left–right” and “global–local.” Higher user activity is associated with more simplified and polarized attitude structures, while highly visible users exhibit attitudes that closely approximate those of the general public, substantially narrowing the gap between social media discourse and survey-based estimates of public opinion.
This work addresses the weak generalization of language models on lexical long-tail (i.e., rare words). We introduce LongTail-Swap, the first zero-shot evaluation benchmark explicitly designed for the long-tail distribution of pretraining corpora. Built upon the BabyLM dataset, it constructs grammatical/ungrammatical sentence pairs centered on extremely low-frequency tokens; model competence in semantics and syntax for such tokens is assessed via zero-shot average log-probability scores over sentence pairs. Unlike conventional benchmarks emphasizing high-frequency vocabulary, LongTail-Swap systematically exposes severe performance bottlenecks on rare words—revealing that architectural differences yield substantially larger performance gaps in the long tail than in the head. Empirical evaluation across 16 BabyLM models confirms the benchmark’s validity and diagnostic utility for probing long-tail generalization.
研究使用儿童式语言模型通过强化学习测试不同形式的反馈对语法学习的支持作用,发现结构对齐反馈最能提升语法正确性。
研究解决了无监督语音模型对不同口音适应的问题,通过引入ABX-Accent基准和使用自适应域归一化方法微调预训练模型来改善跨说话者ABX得分。
研究使用检索增强语言模型测试情景记忆机制是否能缩小词汇频率对句法对比敏感度的影响,发现该方法有效但未完全解决问题。
This study investigates whether political polarization observed on social media accurately reflects the broader public’s attitudes, focusing on the ideological structure in France within a context of issue misalignment. By integrating large-scale X/Twitter data with nationally representative survey responses and employing dimensionality reduction alongside hierarchical modeling, the research examines how structural factors—such as user activity and visibility—influence the expression of political attitudes. The findings reveal that both online and offline political orientations consistently align along two stable dimensions: “left–right” and “global–local.” Higher user activity is associated with more simplified and polarized attitude structures, while highly visible users exhibit attitudes that closely approximate those of the general public, substantially narrowing the gap between social media discourse and survey-based estimates of public opinion.
This work addresses the weak generalization of language models on lexical long-tail (i.e., rare words). We introduce LongTail-Swap, the first zero-shot evaluation benchmark explicitly designed for the long-tail distribution of pretraining corpora. Built upon the BabyLM dataset, it constructs grammatical/ungrammatical sentence pairs centered on extremely low-frequency tokens; model competence in semantics and syntax for such tokens is assessed via zero-shot average log-probability scores over sentence pairs. Unlike conventional benchmarks emphasizing high-frequency vocabulary, LongTail-Swap systematically exposes severe performance bottlenecks on rare words—revealing that architectural differences yield substantially larger performance gaps in the long tail than in the head. Empirical evaluation across 16 BabyLM models confirms the benchmark’s validity and diagnostic utility for probing long-tail generalization.