Small Models Scout Bottleneck Order for Large-Model Data Control

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of indeterminate skill bottleneck resolution order and inefficient data mixing in large language model training. We propose a staged training framework integrating small proxy model exploration with a LogFloor closed-loop controller. This approach transforms bottleneck resolution trajectories into transferable curriculum learning structures, employing a "small-model reconnaissance and path transfer" mechanism to guide large models through sequential bottleneck breakthroughs. Experiments on Qwen2.5 demonstrate that this strategy reduces training tokens by an average of 56.2% and achieves approximately 39% computational savings through cross-scale transfer. These results indicate significant improvements in data efficiency for skill acquisition in large language models, offering a scalable solution to optimize training dynamics and resource utilization.
📝 Abstract
Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether their training trajectories reveal another transferable structure: the order in which larger models should resolve skill bottlenecks. We formulate first-passage skill training, where each monitored skill has a target floor and the objective is to minimize the tokens required to reach all floors. We introduce LogFloor, a closed-loop controller that directs each round toward current bottlenecks, producing phase-ordered resolution trajectories. Across five bAbI skill slices on Qwen2.5-1.5B, LogFloor reduces token cost by 56.2% on average. In 70M-to-12B transfer, three-round replay of a 70M scout path reaches every floor in all eight target runs, saving 30.9% by pair mean, 39.4% in pooled training tokens, and 37.6% under source-cost accounting. On MMLU-control, a frozen scout path succeeds across all eight 12B runs. Collapsing a path to its static marginal mixture or reversing its phase order removes most benefits, while bottleneck labels alone remain partially useful. These results identify phase-ordered bottleneck resolution as a transferable curriculum structure for monitored skill-targeted training.
Problem

Research questions and friction points this paper is trying to address.

Data Control
Skill Bottleneck
Proxy Model
Training Curriculum
Token Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

LogFloor
Bottleneck Order
Small Proxy Models
First-passage Skill Training
Transferable Curriculum
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Seungmin Choi
Stanford University
J
Jiwon Sung
Stanford University
Muhammad Umer
Muhammad Umer
Stanford University
wireless communicationscommunication theory6Gmachine learning
A
Abhiram Rao Gorle
Stanford University
Guijin Son
Guijin Son
Undergraduate, Yonsei University
Natural Language ProcessingLarge Language Models
Y
Youngjae Yu
Seoul National University
J
John M. Cioffi
Stanford University