Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过三阶段方法改进了BabyLM 2026 Strict-Small模型的数据效率,包括前沿建模、原则发现及基于原则的改进,提高了模型性能。
📝 Abstract
Learning from limited text requires models to use context, generalize to new inputs, and retain useful capabilities. Qiushi Engine conducted a long-horizon, end-to-end autonomous research program on BabyLM 2026 Strict-Small, within 10 million corpus words and 100 million cumulative word presentations. Three stages connected frontier advancement, principle discovery, and principle-guided model improvement. Stage I combined compact restatements, budget reinvestment, and residual incremental learning to build a frontier model. Stage II found that exact repetition and aligned restatement produce different patterns of context use, depending on target relations and prediction windows. In controlled tasks, recovering familiar performance did not ensure that unseen inputs could still use learned computations. These findings support a testable data-efficient learning principle: organize experience around the contextual dependencies needed for prediction; separately design visible information, supervision, and preservation; test learning, generalization, and retention. Stage III retained source text, masked more local clues, supervised selected targets, and preserved predictions on ordinarily masked inputs. Two continuation seeds from the same parent outperformed ordinary continuation on the complete nine-metric aggregate. Overall rose from 42.02 to 42.25 across two generations; the second achieved the highest Overall in the public Strict-Small snapshot of 8 September 2026. Further studies addressed compression, relational anchors, shared representations, and measurement. Models are available on Hugging Face; code and research records accompany the GitHub repository. Together, these stages illustrate Research RSI: recursive self-improvement of the research process. Scientific understanding and method innovations change subsequent questions and designs; new experiments test and refine them.
Problem

Research questions and friction points this paper is trying to address.

Data-Efficient
Language Modeling
Limited Text
Context Use
Generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

data-efficient learning
contextual dependencies
recursive self-improvement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shuxing Yang
College of Information Science and Electronic Engineering, Zhejiang University
K
Kaihao Zhu
College of Information Science and Electronic Engineering, Zhejiang University
J
Junjie Yang
College of Information Science and Electronic Engineering, Zhejiang University
R
Rui Zhao
College of Information Science and Electronic Engineering, Zhejiang University
J
Junyao Wu
College of Information Science and Electronic Engineering, Zhejiang University
Y
Yize Wang
College of Information Science and Electronic Engineering, Zhejiang University
Wenhao Li
Wenhao Li
College of Information Science and Electronic Engineering, Zhejiang University
F
Fujia Chen
College of Information Science and Electronic Engineering, Zhejiang University
T
Taowen Deng
College of Information Science and Electronic Engineering, Zhejiang University
S
Shenzhan Hong
College of Information Science and Electronic Engineering, Zhejiang University
Y
Yaqi Li
College of Information Science and Electronic Engineering, Zhejiang University
Z
Zichen Li
College of Information Science and Electronic Engineering, Zhejiang University
J
Jincheng Mi
College of Information Science and Electronic Engineering, Zhejiang University
Y
Yuang Pan
College of Information Science and Electronic Engineering, Zhejiang University
Hongsheng Chen
Hongsheng Chen
Professor of Electromagnetics Academy, Zhejiang University
metamaterialscloaktransformation opticsgraphene
Y
Yihao Yang
College of Information Science and Electronic Engineering, Zhejiang University